2026-09-30 · 6 min read
To build an AI agent that does not make things up, you ground every answer in data you control, require the agent to cite the exact source for each claim, let it fetch facts through tool calls instead of recalling them, give it explicit permission to say "I don't know", and measure all of this against a test set of real questions before and after every change. No single technique removes made-up answers. Together they make errors rare, visible and easy to correct.
This post explains each technique and how we apply them at Syntora Ai when we build AI systems for clients.
Why do AI agents make things up?
A language model generates the most likely continuation of the text it was given. When the answer is in its context, it usually uses it. When the answer is missing, it still produces something fluent and plausible, because that is what it was trained to do. The model has no built-in sense of the difference between knowing and guessing.
Agents add more ways to go wrong. They take several steps, call tools and pass results between steps, so a small error early on can turn into a confident wrong action later. The fixes below target both problems: they keep the model close to real data, and they make every step checkable.
What is grounding and how does it work?
Grounding means the agent answers from specific material you supply at the moment of the question, instead of from what the model absorbed in training. The common pattern is retrieval:
- Your documents, records or knowledge base are split into passages and indexed.
- When a question arrives, the system searches for the most relevant passages.
- Those passages go into the model's context, each with an identifier.
- The model is instructed to answer only from them.
Retrieval quality matters more than model choice. If the search returns the wrong passages, even the best model will answer from the wrong material. Good grounding usually needs:
- Sensible chunking that keeps related information together, such as a full clause of a contract or a full row of a price list
- Hybrid search that combines keyword matching with semantic search, so exact terms like product codes are not lost
- Metadata filters for date, product, region or customer, so the agent does not answer from an outdated policy
- Fresh data, with a known date for when each source was last updated
The same idea runs through our data work. In SyntoraData, every company fact keeps its source page or register and the date it was seen. An agent built on data like that can always say where an answer came from. Our post on B2B data provenance explains why we treat sources as part of the data.
How do citations reduce hallucinations?
Requiring a citation for every claim changes the task. The model can no longer just write a good-sounding answer. It has to point at the passage that supports each statement. That has three effects:
- The model makes fewer unsupported claims, because there is nothing to cite for them.
- Unsupported claims become detectable. A simple check confirms that each cited identifier exists and that the quoted text appears in it. Answers that fail are blocked or sent for review.
- Users can verify answers themselves by opening the source, which builds trust that is earned.
A practical format is to have the model return structured output: each sentence or field paired with the identifier of its source passage. Plain text with footnotes is harder to check automatically.
Why should agents use tool calls instead of memory?
Anything that changes or must be exact should come from a tool, never from the model's memory. Stock levels, order status, prices, account balances, dates and calculations all belong in tools:
- A database query for records
- An API call for live systems such as your CRM or ERP
- A calculator or code execution for arithmetic
- A search tool for documents
Each tool should have a narrow, well-described purpose and return structured data. Log every call with its inputs and outputs, so that when an answer is wrong you can see whether the model misread a correct result or the tool returned bad data.
For tools that change things, such as sending an email, issuing a refund or updating a record, add permission levels. Read-only tools can run freely. Tools that write should either be limited to safe actions or wait for a person to approve.
How do you make an agent refuse when it does not know?
Models lean towards answering. You have to make refusal an acceptable and expected outcome:
- State it in the instructions. If the sources do not contain the answer, say so and offer to pass the question to a person.
- Give it a structured way out, such as an
insufficient_informationresult, so refusing is as easy as answering. - Score refusals as correct in your test set when the answer is not in the sources.
- Set a confidence gate. When retrieval scores are low or the model's own checks disagree, route the item to a review queue instead of answering.
Every system we ship has this gate. Below the threshold you set, a person decides, and their correction becomes a new test case. Our post on AI automation for small businesses in Pakistan shows how that loop works in day-to-day operations.
How do you test an AI agent before launch?
You cannot judge an agent by trying a few questions and liking the answers. You need an evaluation set:
- Collect real questions or documents from the actual workload, including awkward and ambiguous ones.
- Write the correct answer for each, including "not answerable from the sources" where that is true.
- Score automatically where you can: exact field matches for extraction, citation checks for answers, correct tool and arguments for tool calls.
- Use human or model-assisted grading for open-ended answers, with a written rubric.
- Run the full set on every change to prompts, models, retrieval settings or tools.
Track a few numbers over time: accuracy, the rate of unsupported claims, the refusal rate on answerable questions, and the refusal rate on unanswerable ones. A change that improves one and damages another shows up immediately.
The evaluation set becomes the regression test for the life of the system. Corrections from the review queue feed into it, so it grows to reflect the cases that actually cause trouble.
How does Syntora Ai build agents?
We follow the same sequence on every AI project:
- Find the judgement call. A model is worth adding only where a person currently makes a repetitive, rule-following decision. Elsewhere, ordinary code is more reliable, and we say so.
- Build the test set first, from your real data, before anything goes to production.
- Ground and cite. Answers come from your sources, with the source kept on every output.
- Gate on confidence, with a human review queue below the threshold.
- Measure cost and accuracy in production as well as at launch.
Where data cannot leave your network, we build the same thing on open-weight models running on hardware you control, as described on our private AI page.
Working with Syntora Ai
If you have an AI pilot that gives confident wrong answers, or a task you want to automate but cannot afford to get wrong, write to hello@syntorahq.ai or use the contact page. We start with a free 30-minute process audit and will tell you plainly whether a model belongs in the job at all.