ThinnestAI is now an Official Meta Tech Provider
Back to Blog

Keeping a Voice Agent On Script, Without a Flow Editor

T
Thinnest AI Team
Jun 19, 2026• 7 min read
Keeping a Voice Agent On Script, Without a Flow Editor
On Script, No Canvas

Say the absent thing first

ThinnestAI has no visual flow builder. No canvas, no nodes, no edges, no drag-to-connect state machine, and no engine that walks a graph turn by turn. If you came here looking for one, that is the answer and you can stop reading.

The question underneath is a good one, though, and it is not going away: a purely improvising agent eventually improvises the wrong thing. It invents a refund window. It skips the disclosure. It answers a pricing question with a number nobody at your company has ever quoted. On a chat widget that is embarrassing. On a phone call it is spoken aloud, to a person, and it cannot be edited afterwards.

So here is what you actually have to constrain a call, in rough order of how much force each one carries:

  • Withholding tools — the only mechanism here that is a guarantee rather than a strong tendency.
  • Knowledge grounding — most improvisation is a missing fact, not a discipline problem.
  • The system prompt — the script, and an instruction rather than an instruction set.
  • A pinned reply language — removes a whole category of mid-call drift.
  • Escalation triggers — the defined place where the script ends.
  • Call bounds — a nudge on silence, and a maximum duration.

1. Withholding tools — the only real guarantee on this list

Start here, because it is the strongest and people usually reach for it last.

Every tool an agent discovers arrives switched off, enabling one is a deliberate act per tool, and an agent stops at eight. Turning a capability off does not add a line to the prompt asking the model to refrain — it removes the tool. There is nothing to call, so the behaviour is impossible rather than discouraged, regardless of what the caller says or how long the conversation runs.

This is the one mechanism here that is a guarantee rather than a strong tendency. If there is something your agent must never do, the reliable way to enforce that is to not give it the ability, not to write "never do this" in capital letters.

The eight-tool cap is part of staying on script too. Every tool schema is re-sent on every turn, and models get measurably worse at choosing as the list grows. Most agents that wander are agents with too many things to choose between.

2. Knowledge grounding — answer from your material, not from memory

The invented refund policy is not a discipline problem. It is the model reaching for general knowledge because your specific knowledge was not there.

Give it your material: a website crawl one to three levels deep, a sitemap import, pasted text, or file uploads — txt, md, csv, and the text layer of a PDF. There is no OCR, so a scanned policy document is invisible; retype it or paste it as text. CSV chunking is row-aware, which matters because a price separated from its product produces a confident wrong answer, and on a call a confident wrong answer is the expensive kind.

Your own material is always searched before the built-in web search runs. And one knowledge base serves all 37 reply languages, so a Hindi caller and an English caller cannot be given two versions of a policy that quietly drifted apart.

The single highest-yield thing most people can do to reduce improvisation is to spend an hour on the knowledge base rather than an hour on the prompt.

3. The system prompt — the script itself

Tone, boundaries, what to ask and in what order, what to refuse, what to say when it does not know. Writing it for voice is different from writing it for chat: sentences get spoken, so they need to be short, and a numbered list the model reads aloud is a bad experience for whoever is holding the phone.

Be honest about what a prompt is. It is the strongest instruction available and it is still an instruction. A model following a prompt is complying, not executing. That is why the tool section above comes first, and why anything that genuinely must not happen should be enforced by absence rather than by wording.

Change one thing at a time and read the next twenty transcripts. A prompt edited three ways at once teaches you nothing about which edit worked.

4. Pin the reply language

A caller opens in Hinglish and the agent answers in formal Hindi, or drifts into English halfway through. The reply language setting decides which of 37 languages the agent answers in — including all 22 scheduled Indian languages, each in its own script — rather than leaving it to be re-decided every turn from whatever the last sentence sounded like.

This is a small setting that removes a whole category of on-call drift, and it is the one people most often leave on autopilot and then blame the model for.

5. Escalation triggers — the defined exit

Staying on script includes knowing where the script ends. Triggers are chosen per agent: hand over when the agent cannot answer, hand over when the caller asks for a person or is upset, or neither. Off removes the tool, the same way everything else here does.

An agent that hands over at the right moment beats an agent that answers more questions. The reasoning it writes about why it escalated lands in an internal note the customer never sees, and a week of those notes is a precise list of what your script does not cover, in your customers' own words.

6. Bound the call itself

A call that has stopped being a conversation should not stay open. A caller who has gone silent gets a nudge, and then the call ends rather than billing dead air. You can set a maximum call duration, so a stuck call has a ceiling. Neither of these is glamorous and both save money on the day something goes wrong at three in the morning.

What a stranger cannot do to your script

Everything a tenant adds — a crawled page, an uploaded PDF, a connected tool's own description — is fenced and labelled as data before it reaches the model, so ingested text cannot issue instructions to your agent. That boundary is tested rather than asserted: eight live prompt-injection attacks run against the public endpoint as part of the test suite.

The guarantee you do not get

None of the six mechanisms above guarantees the order of a conversation. There is no engine that will refuse to discuss a debt until a disclosure node has fired, or that makes it structurally impossible to ask for a slot before a name. The model is strongly steered by the prompt and it usually complies. It is not executing a state machine, because there is no state machine.

So if your requirement is a machine-enforced sequence — a regulator wants proof that step A always preceded step B, and a prompt that reliably does it is not good enough — this product does not meet it today, and no amount of prompt engineering here will close that gap. That is a real limitation and we would rather write it on this page than have you find it in an audit.

What we will not do is draw a canvas and imply the graph is enforced when the model still chooses the words. A picture of determinism is the most expensive kind of wrong.

How to tell whether it is working

Read transcripts. Every call has one, and it is part of the conversation's own message history rather than a separate store. Read the escalation notes. Read the disposition breakdown before the connect rate — most first campaigns are fixed by changing the calling hours, not the script. Then change one thing.

Test in the browser before a number is involved: the Playground runs the agent's actual configuration, and an in-browser voice call lets you hear the drift you are trying to fix.

Get started

Take one tool away, put one missing document into the knowledge base, and pin the reply language. Then run twenty calls and read all twenty. That is a smaller ritual than building a flowchart and it fixes more.

Build a voice agent free →

No credit card required • Trial includes 25 voice minutes and 200 chat replies

Frequently Asked Questions

Subscribe to our newsletter

Get the latest AI updates delivered directly to your inbox.