Why our voice agent isn't allowed to say "sent"
You can talk to Prio. You call it from the browser, ask what is on your plate, have a draft read to you, add a task while you walk to the train. Voice is the most natural way to delegate, and it is where we learned the hardest lesson of the last few months.
On our own calls, the agent said things that were not true. It said "approved and sent" about an email it had no way to approve or send. It said "draft deleted" when no deletion had run. It said "nothing on your calendar Thursday" without looking at the calendar.
None of these were dramatic. Nothing went out that shouldn't have. But each one broke the only promise a voice agent can make: what I tell you is what happened.
Why voice is different
In writing, an agent's mistakes are visible. You see the draft, the recipients, the card waiting for approval. If the text says "sent" and the queue still shows the email, you notice.
On a call you see nothing. You hear one sentence and you move on. If the agent says "done", you believe it, because there is nothing else to check. A wrong "done" on a call is not a typo. It is a false record of your day, and you only discover it when someone asks why you never replied.
That makes the bar for voice higher than for text: the agent may only claim what it can prove.
Why better instructions didn't work
Our first fix was the obvious one. We told the model, in plain words, never to claim an action it had not taken. We added the rule to the tool results. We added examples.
It did not hold. Voice needs a fast model, because every half second of thinking is half a second of silence on the line. Fast models are good at sounding helpful and less good at following a rule that fights that instinct. In our tests, the instruction "tell the caller that approving happens in Prio" held in roughly one call out of four.
So we stopped asking the model to be honest and started checking.
Five checks that sit between the model and your ear
Every one of these is code that runs on each turn of the call. None of them depends on the model remembering an instruction.
1. A claim needs a receipt
Before a sentence is spoken, a filter reads it. If it says something was sent, approved, deleted or done, the filter checks whether a tool that actually changed something ran on this call. A look-up does not count. Reading your inbox is not sending an email.
If there is no receipt, the sentence is replaced before you hear it. The same filter removes anything the model wrote that was never meant to be heard, like a tool call spelled out as text.
2. Some answers don't come from the model at all
When you say "approve it" or "send it" on a general call, you get one fixed sentence, written by us and not generated: approving happens in Prio, with the full draft in front of you.
That is a product decision, not a limitation we are working around. Approving an email puts your name under it. You should give that signature with the text in view, not on the strength of a summary read aloud while you cross the street. On a call you can still discard a draft you don't want.
3. A question about your day forces a look-up
"Do I have anything Thursday?" is a question about data, not about language. When we detect one, the model is required to call the calendar or inbox tool before it answers. It cannot answer from what sounds likely.
4. The call remembers what it did
Voice platforms pass the conversation back as spoken text only. The model hears what it said earlier, not what it actually did. So it would forget a draft it had already read out, or try to do the same thing twice.
Each turn now carries a short record of the actions taken earlier on the call, built from our own log rather than from the transcript. A named draft beats "that draft". Earlier actions count as receipts for later sentences.
5. It speaks like a person, not a database
Nobody wants to hear an email address spelled out. Recipients are spoken as names. Formatting that only makes sense on a screen is stripped before the text reaches the voice, and our voice tests are graded on what the caller actually hears, not on the text the model produced.
What this cost us
Some speed, for one. A check on every sentence adds milliseconds, and a forced look-up adds a round trip. We spent a good part of September winning that time back elsewhere: running the voice route next to the database, warming the model's cache when a call starts, and having the agent say a short "one moment" when a look-up takes longer.
It also made the agent less smooth. A model that is allowed to improvise sounds friendlier. One that has to prove its claims sometimes says "I can't do that on a call, it's waiting in Prio." We think that is the right trade. An assistant you have to double-check is not saving you time.
What to ask any voice agent
If you are evaluating a voice assistant for real work, three questions separate a demo from something you can rely on:
- When it says "done", what checks that it's true? If the answer is "the model is instructed to be accurate", that is a hope, not a check.
- Can it act on something you haven't seen? Sending an email on the strength of a spoken summary is a risk you take on, not the vendor.
- Does it look before it answers? Ask it about tomorrow's calendar and then check whether it actually read it.
We wrote earlier about the trust ladder AI agents should climb. Voice doesn't get to skip a rung because it sounds human. If anything, it has to earn each one twice.
Voice calls with Prio are part of the Pro plan. Try it on your own day: ask what needs you, have a draft read out, and listen for the moment it tells you the approval is waiting in Prio.