An answer with the instructions showing
The tool answered questions about official documents. For each question, the assistant received passages from several sources in priority order and had to explain the subject to the reader. On screen, though, part of the answer began discussing layers of authority, as if the reader needed to know how the context had been arranged behind the scenes.
This was more than a matter of style. When the final text repeats labels used to organize the material given to the model, it becomes harder to tell what came from a source and what came from its packaging. The answer can sound methodical while saying less about the document the person wanted to understand.
The ordering made sense. The backstage vocabulary chosen to describe it did not.
A heading joins the conversation too
In a system that assembles documents for an answer, context construction is not neutral. Headings and separators enter the model's view alongside retrieved passages. Call a section “highest authority source” and that phrase may return in the synthesis, even if the user never asked for it.
A prompt instruction saying “do not talk about layers” might seem protective. But it would leave the cause in place: the context would keep repeating the phrase. Changing what the model received was more robust than hoping it would ignore that label every time.
Priority did not need to go away. The blocks kept their order; their headings began naming the sources instead of narrating the internal mechanism that ordered them.
Two vocabularies for two audiences
The system still needed machine-readable tier IDs to carry source metadata through the API and preserve priority. Those IDs served the code. The reader needed a document name and a passage that supported the answer.
Context assembly began using source titles in the text shown to the model while retaining technical IDs in structured data. When retrieval returned an empty source, the context stated the gap explicitly rather than leaving a space inviting the model to complete the answer on its own. That improved how the context read, but depended on retrieval completing without error: a retrieval failure could also produce an empty list.
It is a small separation in code and a large one in reading: metadata for the machine, source names for the person.
What the test actually showed
Verification covered context assembly: priority order was preserved, the old internal headings no longer appeared in the text sent to the model, and empty lists were signaled. The API's machine-readable IDs remained stable. Those tests did not distinguish a search with no hits from one that failed before consulting the source.
That test does not prove a model will never repeat backstage language again. It proves a narrower point: this route for leaking it, created by how we labeled context sections, was removed from the input. Evaluating the complete answer still requires observing real outputs, not only the template that feeds them.