How to make an AI assistant reliable and keep hallucinations down

Abylkaidarov M.K., lawyer, beginner AI engineer, ND SaaS Solutions LLP

How it started

I am a practising lawyer, and I decided to build an AI assistant for the legislation of Kazakhstan — by today that has taken about two months. The assistant has one job: find the norm in the regulatory legal acts (hereinafter — RLA) and name the article, paragraph and sub-paragraph that answer the question.

First I wrote down, in full, the way I work: how I look for the answer to a question. It came to eight steps — parsing the question itself, searching for directly and indirectly relevant acts, selecting the applicable norm, reasoning, and composing the answer. On that basis I wrote the "search and selection" standard, which became the foundation of the search-and-selection nodes and of the way acts are prepared for the vector database (more on that below).

Second, I have always read the articles of lawyers I respect — professors of the Kazakh State Law Academy, where I studied in 2000–2004 — in order to follow their line of thought and apply it in my own work. I wrote out my own reasoning, collected the publicly available articles of Didenko A.G., Suleimenov M.K., Basin Yu.G. and Karagusov F.S., and drew rules from them. Together with mine that came to about thirty; fifteen went into the "Lawyer" standard — how to find a norm and how to answer. Not everything could be carried over into the AI: some reasoning is peculiar to a human being, and some of it produced hallucinations — the AI started taking information from outside the vector database (analogy of law and analogy of statute, for example). Everything that calls for legal judgement — filling gaps from general principles, evaluative notions such as "reasonable" and "in good faith" — was rejected deliberately. Instead the assistant says honestly that it has not found an applicable norm, or has found norms for only part of the question.

The "Lawyer" standard became the foundation of the search-and-selection nodes and of the main model — the assistant's "brain".

Building the assistant

1. I prepared the product concept with the help of three AIs (GPT, Gemini, Claude). Each contributed something: the choice of tools, finding documentation, testing hypotheses.

2. Development ran in three directions, sometimes at once, sometimes one after another — depending on what turned out to be the bottleneck.

2.1. Preparing the acts. You cannot load acts into a vector database as they are (docx, pdf, html): acts carry a great deal of "noise" that gets in the way of search — subordinate acts especially, where the structure is harder to determine. Each kind of act has its own cleaning program: it strips the noise — headers, signatures, service notes, lists of amendments — cuts the act into articles and paragraphs, and puts a precise label on every piece: act, version, article, paragraph. It is that label which later lets the assistant cite the norm and give a link to the primary source. At the time of publication, in September 2026, the database held 3,791 documents: 23 codes, 16 constitutional laws, 248 laws, 205 normative resolutions of the Supreme Court, the Constitutional Court and the Constitutional Council, 104 presidential decrees and 3,195 subordinate acts — about 425 million characters in all. A separate set covers the law of the Astana International Financial Centre: 54 AIFC Regulations and Rules, 10 Kazakh acts on the Centre and 7 AIX documents; those texts are in English, and the answer comes in Russian with the article cited and a link to it on the AFSA portal.

2.2. The search-and-selection nodes. These are blocks of code (Python) and "light" AI models with short prompts for specific tasks. Their job is preparatory: find the applicable norms in the vector database and hand them to the main model.

2.3. The main model — the assistant's "brain". A heavy model, able to hold a complex prompt together with the incoming norms. The prompt is some forty numbered rules: the "Lawyer" standard, the reasoning logic, the structure of the answer, and the constraints that keep hallucinations down.

2.4. Delivering the answer. The formatting of the answer and the links to the primary source are built by a separate layer, not by the model. The model is not responsible for presentation, and that too reduces errors: the fewer jobs one prompt has, the less often it confuses them.

All told the assistant is 39 nodes and 43 connections, 9 of them calls to AI models.

How I worked on reliability

1. I count an answer reliable if: every norm it names exists in the database and its paraphrase matches the text of the act — the number, the term, the persons covered, the condition; the address of the norm is exact — act, version, article, paragraph; the answer addresses the question asked rather than a neighbouring one; it contains nothing beyond the norms found — no "general principles", no analogies; and if the norm is not in the database, the assistant says so.

2. Answers are measured by a digital judge, and I have two of them, from different developers, measuring in parallel. A digital judge is a set of code plus a "light" model that checks the answer on the merits and for invention, and verifies every norm it names against the text of the act in the database. Setting the judge up took a long time: the early versions erred towards "the assistant is bad" — they truncated the answer, they counted correct behaviour as a defect. On the very same answers only the judge changed — and the score moved by half. Until I was satisfied that the judge could be trusted, its readings were not used.

3. A single judge is systematically blind in its own direction. On 111 saved answers one judge passes 94–96, the other 75–79. Reviewing the disagreements by hand, with the text of the norm in front of me, showed: three real errors by the assistant, thirty-two false accusations. One lets things through, the other over-reports, and on its own each gives a false picture. So what goes into the work is not the score but the list of disagreements between the two judges: an hour of that review yields more real defects than a day spent with a single judge.

4. I learned to measure which stage got it wrong. Checking every norm named against the text of the act showed: in most of the misses the correct norm was already in front of the model, and the model misstated it — the wrong number, the wrong term, the content of the neighbouring article, an extended reading. Only in a minority of cases was the norm missing and the model named it from memory. Until that measurement I had been fixing search; after it, it became clear that the main loss is in the answering model.

5. Guided by the judges and by my own control checks, reliability grew: it started at 50 % and reached a band of 82–86 % of sound answers (84 % today, 93 out of 111). The work continues; the target is 90–95 %.

In the end

Over two months I went through five stages of accepting AI:

1. Admiration — the AI did things I did not know and could not have known how to do: it set up a server, wrote code, tested hypotheses very quickly, processed large volumes of data.

2. Distrust — the AI lied plausibly and invented things across the board, and admitted it honestly; I had to ask the same question several times until I began to see how it had reached its conclusion.

3. Disappointment — once legal questions began, the AI managed to add things of its own even under a direct instruction to answer only from the vector database.

4. Understanding — it did not come at once, only through repeated testing of hypotheses and measurement of results. I learned to measure which stage exactly made the error: the search-and-selection nodes or the main model.

5. Acceptance — I started putting my requests to the AI properly, in full and with structure, stating my expectations, and measuring the result with verified instruments against thresholds set in advance.

Can AI be trusted? It can, but not blindly