e10 Infotech - AI-powered software development

RAG development that answers from your documents and shows its working

Retrieval augmented generation is mostly a search problem wearing an AI hat. If the right passage never reaches the model, no amount of prompting rescues the answer. e10 Infotech Private Limited builds retrieval systems the way search engines are built: measured recall, hybrid matching, reranking, and evaluation against questions your staff actually ask. Answers carry citations that resolve to the source paragraph, and the system refuses when the evidence is not there. Businesses in New York City come to us when staff or customers need reliable answers out of documents nobody has time to read.

What we build

  1. Internal knowledge assistants over policies, manuals, tickets and wikis
  2. Customer facing answer systems grounded in documentation and product data
  3. Contract, tender and case file analysis with clause level citation
  4. Research and literature tools across large document collections
  5. Support deflection systems that answer from history and known resolutions
  6. Code and technical knowledge retrieval across repositories and specifications
  7. Multimodal retrieval covering scanned documents, tables, diagrams and images

How a retrieval system is built

Ingestion is where most quality is won or lost. Documents are parsed with layout awareness so tables, headings and figures survive, chunked on semantic boundaries rather than fixed character counts, and enriched with metadata such as source, date, owner and access level. Retrieval combines dense vectors with keyword search, because names, codes and part numbers fail on embeddings alone, then a reranker orders the candidates. We measure retrieval separately from generation, so when an answer is wrong we know which half caused it. Permissions are applied at query time against your identity provider, so a user never sees a passage they could not open directly. Freshness is handled with incremental re-indexing tied to your source systems rather than a periodic full rebuild.

Where this fits

Retrieval sits inside our AI First Development practice and underpins most of it. Answers are produced through LLM Integration and Fine Tuning and Generative AI Development, delivered conversationally by AI Chatbot Development, and used as a tool by AI Agent Development. Bulk document processing runs through AI Workflow Automation, hosting and vector infrastructure through MLOps and AI Infrastructure, and scoping through AI Consulting and Strategy. Engineering teams pair it with Building with AI Coding Tools, and the surrounding application comes from Software Development and Web Development.

Accuracy, citations and refusal

Every claim in an answer is tied to a retrieved passage, and the citation is checked against that passage rather than accepted because the model produced it. When retrieval returns nothing relevant, the system says so and offers the closest sources instead of composing something plausible. Conflicting sources are surfaced with dates so the reader can judge, which matters when a policy has been superseded but the old version is still in the drive.

Measuring whether it works

We build a question set from real enquiries and score retrieval recall at rank, answer correctness, citation validity and refusal appropriateness. Those numbers gate every release. In production we track unanswered questions, thumbs down responses and queries with weak retrieval scores, and that queue drives the next round of ingestion and tuning.

What you get at handover

Source code in your repository, infrastructure as code in your cloud account, the ingestion pipeline with parsers and chunking configuration, the vector and keyword index definitions, the evaluation question set with scores, permission mapping to your identity provider, tracing dashboards and cost per query, a re-indexing runbook, and a thirty day warranty on behaviour that does not match the specification.

Working with businesses in New York City

Work for clients in New York City runs remote first: a named engineering team, a scope agreed in writing before anything starts, and demos on a fixed cadence you can hold us to. Working hours overlap your business day and everything is delivered in English.

Rules on where documents may be processed and how long derived data may be kept differ by jurisdiction, so we confirm what applies to businesses in New York City before choosing embedding and hosting providers. That answer shapes the architecture rather than following it.

You get one point of contact, senior engineers on delivery, and reporting tied to retrieval and answer scores rather than activity. Businesses in New York City and the wider region are supported on the same terms.

Ready to get answers out of your documents?

Tell us what people keep asking and we will send an approach, an evaluation plan and a quote.
Serving New York City and the wider region.

Talk to e10 Infotech

§QA

Queries raised before signature

Everything worth knowing about RAG and Knowledge Systems in New York City.

01What is retrieval augmented generation?

It is searching your own content for relevant passages and giving them to a language model as context, so the answer is grounded in your material rather than in the model's training data.

02Why not just fine tune the model on our documents?

Fine tuning teaches style and format well and facts badly. It also cannot handle documents that change weekly, and it gives you no citation. Retrieval solves all three.

03How accurate is it?

Accuracy depends on retrieval quality more than model choice. We measure both separately against a question set from your real enquiries, and we report the numbers rather than a demo.

04Does it respect our existing permissions?

Yes. Access is enforced at query time against your identity provider using document level metadata, so a user only ever sees passages they could open directly in the source system.

05What document types can you handle?

Office documents, PDFs including scanned ones, spreadsheets, email, wikis, ticket systems and code repositories. Layout aware parsing keeps tables and headings intact through ingestion.

06How long does a RAG project take?

A grounded assistant over a defined document set usually reaches production in six to ten weeks. Wider rollouts with many sources and strict permissions take longer.

07How do you keep the index current?

Incremental re-indexing triggered by changes in your source systems, with deletions honoured so withdrawn documents stop being cited immediately.

08What happens when the answer is not in our documents?

The system says it cannot answer and shows the nearest sources. Refusal quality is scored in the evaluation set, because a confident wrong answer costs more than no answer.

09Which vector database do you use?

pgvector where you already run Postgres and want one fewer system, Qdrant or Pinecone at larger scale, and hybrid search alongside a keyword index in every case.

10Can it work across languages?

Yes, with multilingual embeddings and reranking, so a question in one language can retrieve and cite a document written in another.

11Why do we need keyword search if we have embeddings?

Because embeddings are poor with exact identifiers such as part numbers, invoice references and surnames. Hybrid retrieval catches both meaning and exact matches.

12What does it cost to run?

Cost is per query plus index storage and re-indexing. Caching, smaller models for extraction and tighter context keep it predictable, and we set a target during design.

13Do you offer RAG development in New York City?

Yes. e10 Infotech delivers RAG and Knowledge Systems for businesses in New York City and the wider region, remote first with a named team and a scope agreed before work starts.

14How do you run RAG projects for clients in New York City?

Working hours overlap the New York City business day, communication is in English, and you get one point of contact with reporting tied to retrieval and answer scores rather than activity.

Enquiry

Talk to an engineer about RAG and Knowledge Systems in New York City.

Four lines is enough to start. You will hear back from someone who does the work, usually the same day, and nothing is sent to a sales desk.

Prefer to set it all out at once? The full brief form asks the longer questions. Or write to [email protected].

Short form shown on service and industry pages. The long qualifying form stays on /contact.