How we built a voice interviewer to capture project knowledge that never reaches the documentation, and the three constraints it took to make it safe inside a production application: user-scoped authorization, an eighteen-tool read-only surface, and tools that report only outcomes the interface actually produced.
01
A voice agent for the documentation nobody writes
02
What a specification does not say
A specification tells you the field exists and what type it is. It rarely tells you why it was added, which two teams argued about it for a month, or that everyone on the project knows you never populate it for German customers because of a decision taken in 2019. That part lives in a hallway conversation, a Slack thread that scrolled away, a remark someone made while sharing their screen.
The usual response is to ask people to write it down. Usually they do not. Writing is slow, and the person holding the knowledge is the person with no time.
People will talk, though, and not only because talking is quicker. Polanyi put it as "we can know more than we can tell" [C31]: much of what an expert knows is tacit, dependable in practice but not in any form they could simply type out. Nonaka's account of how organizations turn that into recorded knowledge treats the step from tacit to explicit as a mode of its own, and dialogue drives it [C32].
That matches what we see. Ask someone to spend ten minutes explaining a project out loud and they will tell you things they would never have written down, including parts they did not know they knew. So we put a voice interviewer inside our product. It asks about the project in the language the project runs in, and what gets said can become a document the rest of the platform uses.
Then you notice what you have built: a model with a microphone inside a production application, talking about confidential work, with tools that reach real data. That is not a documentation feature. It is an authorization problem with a friendly voice.
03
The short version
Three constraints make the interviewer safe to run inside a production application.
It acts as the person speaking to it. Every tool call carries that user's own credentials, whether it runs in their browser or on our server, so the access control we already have does the enforcing.
It sees eighteen tools out of 161. Handing it the full catalog would cost 40,737 tokens before anyone says a word. OpenAI's soft suggestion is fewer than twenty per turn [C16].
It has no tools that change project content. The tool catalog admits only tools declared read-only, and the build fails if someone adds one that is not.
Two more things hold today. The navigation tools report only what the view actually rendered. And with the provider we currently use, the audio reaches neither our backend nor our model gateway [C15].
Writing to the project is not available yet. When it is, it will need the person's confirmation, enforced on the server.
04
What it is for
Someone opens a project in Testuvion and talks. The agent asks what the project is for, what the system does and which rules matter: the things a specification assumes you already know. While it asks, it has the project's documents and processes in front of it, along with the page the person has open and what they last changed. Its questions start from what the project already records, and it can follow up when an answer contradicts what is written down. Spoken and typed turns land in the same thread.
When the person wants to keep the conversation, one action turns the thread into a knowledge document. It lands among the project documents the platform already generates test cases from, so a rule that lived only in someone's head can be traced to a test. A recorded meeting gives you an hour of audio that is rarely opened again. A conversation with a goal, run by something that already knows what is documented, gives you the part that was missing.
The feature rests on one decision: the conversation runs on a speech-to-speech model. Audio goes in, audio comes out, and nothing transcribes your words before the model answers them. The alternative chains speech to text, text to a language model and text back to speech. OpenAI's guide draws the line in the same place: the realtime path is "Best for: Speech, reasoning, and tool use in one session", and the chained one is for when "each stage needs to be visible or replaceable" [C47].
Each hop in the chained path adds delay before the first sound comes back, and the two text stages are where tone, hesitation and interruption stop reaching the model as sound [[C47](https://developers.openai.com/api/docs/guides/voice-agents)]. We take the one-hop path.
For an interview, that matters. In our experience, a relay that pauses before every reply pushes people into dictation: short, complete, careful sentences, the register in which the useful thing rarely gets said. You want someone thinking out loud, wandering, correcting themselves and cutting the agent off when it has misunderstood, and that needs replies fast enough to feel like a conversation.
For the same reason, with an open microphone we use semantic turn detection. It weighs whether you have finished the thought as well as whether you have gone quiet, so a pause while you recall why a decision was made in 2019 need not end your turn. Where an open microphone is impractical, the same conversation runs push-to-talk, and pressing the button cuts the agent off.
Czech brings one more limit. The realtime model speaks Czech with English phonetics, and no prompt we have tried changes that. It can also switch to English partway through a session, an upstream issue reported for several languages. That part responds to prompting: the Czech interviewer runs on a prompt written in Czech, with a positively stated rule that adapting to the speaker's accent never changes the language of the answer. We have not yet measured how often the switch still happens.
Voice also moves you around. Ask for a document and it opens. Ask about a process and it appears in the tree. Ask where a claim came from and the passage scrolls into view, highlighted. In a large application, saying what you want is easier than remembering where it lives.
So it reads and it navigates, and it changes no project content. Keeping it that way took most of the engineering.
05
It acts as the signed-in user
Give an agent access to an application and you must decide whose identity it acts under. The Open Worldwide Application Security Project (OWASP), whose Top 10 risk lists your security team already reads, names one answer as a root cause of Excessive Agency: an integration built to act for one person that reaches the systems behind it "with a generic high-privileged identity", such as a document reader connected "with a privileged account that has access to files belonging to all users" [C20]. Its remedy is to execute actions "in the context of that specific user, and with the minimum privileges necessary" [C21].
A policy like that can drift, so we built it into the structure: the agent has no identity of its own. Every tool call starts in the signed-in user's browser and carries that user's credentials. Three tools act on the screen and run there. The rest go to one endpoint on our backend, which looks the tool up in a catalog, refuses anything the catalog does not name, and runs it with the rights of the person whose credentials came with the request. Those are the same rights the documents page or the process tree would apply to that person. The endpoint asks only for read access, because nothing behind it writes.
In this design the model never receives an API credential of ours. It names a tool and code does the rest, which is OWASP's other instruction: handle these functions in code rather than giving them to the model [C18].
So we never have to decide separately what the agent may see. The access control we already need for signed-in users answers that. There is no agent identity to provision and no second permission model to keep in sync. If you cannot open a document, neither can the voice talking to you about your project.
The audio path is where the claim is easiest to overstate. Our backend mints a short-lived session credential and hands the browser a URL. The browser negotiates its own connection. Our model gateway relays the handshake, and its documentation says that "audio never passes through the proxy" [C15]. The speech reaches neither our application backend nor our gateway. It does reach the model provider, which is what matters if you have a data-residency question. And it holds because of the provider and the relay we currently use, both of which are configuration. Change either, and the answer can change with it.
Two channels leave the browser, and they carry different identities:
The audio goes to the provider, and the gateway relays only the handshake [[C15](https://docs.litellm.ai/docs/proxy/realtime_webrtc)]. Every tool call that leaves the browser returns to our own API as the speaking user, under the same access control as the rest of the product.
06
Eighteen tools out of 161
The obvious way to make a voice agent useful is to give it everything. We already expose our whole backend as a curated catalog over the Model Context Protocol (MCP), an open standard for exposing tools and context to an agent. The catalog holds 161 tools, and putting it on a voice session is a configuration change.
We measured what that would cost. Serialized as a model receives them, the 161 definitions come to 40,737 tokens. The eighteen the voice session ships come to 2,156. The difference is spent on every session, as input, before the first word and before any provider-side caching.
Accuracy is the other objection. OpenAI's guidance is to "aim for fewer than 20 functions available at the start of a turn at any one time, though this is just a soft suggestion", and to keep the number small "for higher accuracy" [C16]. Eighteen sits just under that line. The surface is that size because the voice agent and the text assistant share one catalog, so every read the text assistant needs, the voice agent gets too. We accept that trade for capability, and it leaves less headroom than we would like.
The usual answer to a large catalog is deferral: expose a few tools and fetch the rest on demand, which the same guidance recommends [C16]. We plan that one level up. The voice agent stays small and would hand anything bigger to a server-side agent that holds the full catalog on its behalf, so the deferral happens between two agents. OpenAI's guide describes this shape as its third architecture: a live model that "delegates reasoning and tool use to a separate backend" [C47]. Filtering the tool list per turn inside the session is the other option, and we have not evaluated it yet.
One option we rejected outright: letting the voice session call our tool catalog directly from the provider's infrastructure. The connection would then start outside the browser. Our endpoint would have to be reachable from the provider, and the user's credentials would sit in provider-side session state. That turns an authorization design into an exfiltration surface, so the credentials stay in the browser.
07
A tool that cannot report what it did not do
An agent will tell you it did something, and voice removes the usual way of checking. There is no message on screen to compare with reality, only a confident sentence you have already heard. Where the point is a trustworthy record, an agent that cheerfully claims it stored something is worse than no agent.
The tools that change what you are looking at report only what the view confirms. When the agent highlights a passage, it files the request, navigates and waits for the view that actually rendered to answer: found, or not found. A late answer to an older request cannot pass for the current one. If nothing answers within seven seconds, or the request was cleared first, the tool returns "pending", which means nobody observed the outcome, as distinct from a passage nobody found. In Testuvion no path returns success without the document view having set it.
That is the part that carries over to other agents: the honesty lives in the tool's return contract. A system prompt asking a model not to claim what it has not verified is a request. A tool that can only return what the view resolved is a constraint.
The transcript has no such check yet. The speech-to-speech model gives us no text of what you said. The transcript we keep comes from a separate transcription the provider runs alongside the conversation, and since the audio never reaches us, the browser reports it back. Nothing on our side re-derives it. This is exactly what OpenAI lists in favour of the chained path: "you might store the transcript, run policy checks before the text agent responds" [C47]. Checking our transcript would mean comparing what was written with what was said, a harder mechanism than resolving a highlight, and we have not built it. For a knowledge document that a person reviews, a record taken on trust is acceptable. If that record ever had to settle an argument, it would not be.
08
How we approach mutations
The more useful version of this feature books the meeting, creates the test case and updates the document. We have deliberately not opened that half, because the safety argument has to come before the capability.
Today nothing on the voice surface can change project content, and the build enforces it. Every tool the agent can call is declared in one catalog, and the catalog refuses to load if any entry is not marked read-only. A test fails the build if a catalog module so much as mentions a database write, and another fails it if anyone declares a voice tool outside that catalog. The endpoint that runs tools accepts only names the catalog knows. A determined contributor could still label a write as a read, so this is a failing build and not a proof. But the property does not depend on anyone remembering it, and that is what a gate for writing can be built on.
OWASP says that "it is unclear if there are fool-proof methods of prevention" for prompt injection [C17], and pairs its user-context rule with a second one: require human approval for high-impact actions [C21]. Simon Willison puts it harder: once an agent has ingested untrusted input, "it must be constrained so that it is impossible for that input to trigger any consequential actions" [C19]. Not unlikely. Impossible. A capture agent ingests content written by other people all day. That is its job.
So the design sorts every tool into three tiers. Read tools stay allowed outright, as they are today. A middle tier would be able to change things, but only in two turns. The first turn could not call them at all: the agent would plan, resolve what you meant, and come back with a summary and a short-lived signed token. Only a second turn, carrying that token and your confirmation, would execute. A third tier would never be reachable by voice at any privilege: deleting a project, removing a team member, touching integration credentials. Speech is a lossy channel, the microphone may be open, and some operations should not be one sentence away from happening.
The two-phase gate as designed.
What matters is where the enforcement would live. In this design the first turn cannot call a write tool at all. The token it returns is short-lived and bound to the session and to the exact plan in the summary, so a later turn cannot carry out a different action from the one the person heard. A write then needs a second turn with that token, after a confirmation the person can see in the browser. A poisoned document can still make the agent propose something; it cannot skip the confirmation or swap the plan. One question is still open in the design: whether a spoken "yes" is enough, or some writes should always need a tap on an on-screen card.
None of that is open yet. Today the agent has nothing to write with, because the catalog refuses to register a writable tool. When writing arrives, it arrives behind the gate.
09
Three questions for your own agent
None of this depends on voice or on our stack. If you have an agent touching a real application, you can answer three questions this afternoon.
What identity does it act under? Look at whose credentials are in the request when it acts. If it is a service account, your blast radius is whatever that account can reach. Where that exceeds the person in front of it, so does the agent, and no prompt changes that.
If it were talked into something, what would stop it? Answer with a mechanism. "The instructions tell it not to" does not count: instructions are a request, and prompt injection exists because requests can be outargued. Look for the point where a compromised model's intention cannot be carried through: a permission it does not hold, a gate it cannot open, a signature it cannot produce. If the only protection is an instruction in the prompt, nothing technical prevents the action.
Can any tool report a success it did not achieve? Find one that changes state or the user's view and read its return path. If there is a way back that does not consult reality, the agent will eventually take it.
For us, the first is a property of the design, and the third holds where we thought to ask. The second has two answers. For reading, a catalog that rejects anything not declared read-only, backed by checks that fail the build. For writing, a gate we have designed and not yet built, which is why the half of this feature that changes things is not open.
The knowledge that never gets written down is worth capturing, because quality assurance decides from it anyway. A rule that lives only in someone's head cannot be traced to a test, challenged by a reviewer or rechecked when the system changes. Voice is another way to get that rule into the system, and it adds an obligation to verify what you collected. The thing you built to collect it is a model with a microphone inside your production system, and it deserves the same scrutiny as anything else that fits that description.
10
Models used, and how this was written
The interviewer speaks through a hosted realtime speech model, selected in configuration rather than hard-coded. The knowledge document extracted from a conversation is produced by Claude Sonnet. Neither model sees anything the speaking user could not already open. The illustration at the top of the article was generated with an OpenAI GPT image model.
The token measurements were taken on 17 September 2026 against a pinned snapshot of the catalog with the same 161 tools as the live one, and a voice surface of eighteen. They move whenever either does and are due to be re-run in March 2027. They describe that configuration, not a fixed cost.
This text was drafted with AI assistance from a list of checked sources, and edited and approved by its author. Every number in it traces to one of them: measurements to a command we ran, external claims to a published work.
CTA
See it run on your own specification
We’ll run a narrowed part of it live on the call and walk through the output together, and you’ll have the complete generated suite within 24 hours.
Visual testing checks a built page against its Figma frame, text node by text node, and gives coding agents a pass or fail instead of a screenshot to squint at.
Human oversight of AI mostly fails, and the research says why: accountability alone does not make people check, and showing the reasoning is contested. What is left is the cost of verifying and whether the decision can be skipped. This is the case for treating a tester's approval as the moment a generated test case acquires standing, not as a safety net bolted on at the end.
[C19] Simon Willison, "The lethal trifecta for AI agents: private data, untrusted content, and external communication", 16 June 2025. simonwillison.net