⚡ Making full use of Flux eager end of turn detection with Jev
Jev Runner is a 10x10 maze you steer by voice while other people in the room keep talking.
Every voice agent has to answer two questions on every turn, and turn detection only answers the first: when did the user stop talking, and what did they want? Most builds answer them in sequence. They wait for the turn to end, send the transcript to a model, and then wait again for a response.
Deepgram Flux has an event built to overlap those two waits. EagerEndOfTurn fires when Flux thinks a turn is probably ending, before it's sure, and the Flux launch post says what it's for: "calling an LLM speculatively, i.e., in preparation for an upcoming turn end, in order to minimize response latency."
Cost has always been the barrier. A speculative call that turns out wrong gets thrown away, and a thrown-away LLM call still costs tokens. TypeSafe's Jev changes that math. As of September 2026, TypeSafe prices Jev at "$0.042 per one million input tokens, while output tokens are completely free," and its launch post calls output "too cheap to meter." The call in this post uses about 3,800 input tokens, so each one costs about $0.00016, or 0.016 of a cent. At that price you can speculate on every eager event and stop worrying about the ones you discard.
Key takeaways
- Flux's
EagerEndOfTurnevent is documented for speculative model calls. Call Jev on it, hold the answer, and commit only onEndOfTurn. - A Jev call in this build costs about $0.00016 and takes 140 ms at the median, so a discarded guess costs almost nothing.
- When the settled transcript matches the eager one, reuse the held answer. In a 50-second session with office chatter, the answer was already waiting on 11 of 11 turns.
- Speculating this way made about three times as many model calls as waiting would. At Jev's price, the whole session cost 1.11 cents.
- Don't act on the eager answer. The words that make room speech irrelevant often arrive last.
- Gate on the probability of an explicit
noneoption, since selected-option confidence drops for reasons unrelated to relevance.

How do I call Jev on Flux's eager end of turn?
Deepgram's guides to Eager End of Turn and Speculative Replies cover the turn detection setup. This post covers pairing it with Jev, and what that costs.
Open Flux with an eager threshold below the end-of-turn threshold. These examples use the Deepgram JavaScript SDK (npm install @deepgram/sdk, tested with 5.13.0 on Node 24), and each one was run against the live Flux and Jev APIs before publishing. The client reads DEEPGRAM_API_KEY from your environment. If you don't have one yet, sign up for a free key.
import { DeepgramClient } from '@deepgram/sdk'
// DeepgramClient reads DEEPGRAM_API_KEY from the environment.
const deepgram = new DeepgramClient()
const flux = await deepgram.listen.v2.connect({
model: 'flux-general-en',
encoding: 'linear16',
sample_rate: 16000,
eot_threshold: 0.7,
eager_eot_threshold: 0.3,
})Then handle the two turn events differently. askJev is defined in the next section, and commit and reportError are whatever your app does with an answer or a failure. The SDK keeps one handler per event, and a second flux.on('message', ...) replaces this one, so keep all of your turn logic in a single handler.
const held = new Map<number, { transcript: string; answer: ReturnType<typeof askJev> }>()
flux.on('message', m => {
// The SDK hands you parsed messages, so there is no JSON.parse here.
if (m.type !== 'TurnInfo') return
const text = (m.transcript ?? '').trim()
if (!text) return
// Eager: ask now, commit nothing. The sentence may still be arriving.
if (m.event === 'EagerEndOfTurn') {
const answer = askJev(text)
answer.catch(() => {}) // a discarded guess should not raise, the settled pass reports errors
held.set(m.turn_index, { transcript: text, answer })
return
}
// Settled: reuse the held answer if the words did not change.
if (m.event === 'EndOfTurn') {
const spec = held.get(m.turn_index)
held.delete(m.turn_index)
const answer = spec && spec.transcript === text ? spec.answer : askJev(text)
answer.then(commit).catch(reportError)
}
})Register the handler, then open the socket and start streaming audio with flux.sendMedia(chunk):
flux.connect()
await flux.waitForOpen()The whole design rests on three rules:
- On
EagerEndOfTurn, call Jev and hold the answer. - On
EndOfTurn, reuse the held answer when the transcript matches, and ask again when it doesn't. - Only
EndOfTurncommits an action.
Flux can fire EagerEndOfTurn more than once in a turn, and it can send TurnResumed when the speaker keeps going. The map keyed on turn_index handles both cases, because a newer eager event overwrites the held answer and a changed transcript at EndOfTurn triggers a fresh call.
Send the Jev request
Jev answers structured questions about a piece of text. You send the text as state along with a set of Choice questions, and each answer comes back with the selected option, a probability for each option, and a confidence score.
async function askJev(utterance: string) {
const res = await fetch('https://api.typesafe.ai/v1/systemone', {
method: 'POST',
headers: {
Authorization: `Bearer ${process.env.TYPESAFE_API_KEY}`,
'Content-Type': 'application/json',
},
body: JSON.stringify({
state: utterance,
model: 'jev-latest',
questions: {
dir_1: {
type: 'choice',
instructions: 'Which direction is the first maze move? Use none if the speaker is not addressing the maze.',
criteria: {
up: 'Move the maze marker up. A bare command like "up" or "go up two" counts.',
down: 'Move the maze marker down. A bare command like "down" counts.',
left: 'Move the maze marker left. A bare command like "go left" counts.',
right: 'Move the maze marker right. A bare command like "go right two" counts.',
none: 'No move. The speaker is talking to another person, or about a document, a webpage, a schedule, a deadline, or a price, even when they use a direction word.',
},
},
},
}),
})
if (!res.ok) throw new Error(`jev ${res.status}: ${await res.text()}`)
return res.json()
}To test all of this, I built Jev Runner, a 10x10 maze you steer with your voice. The version I play asks thirteen questions per turn: how many moves, then a direction and a step count for each of up to six moves. Those questions are where the 3,800 input tokens come from.
How much time does the eager call save?
It depends on the gap between EagerEndOfTurn and EndOfTurn, and that gap varies. I streamed 20 seconds of recorded office audio through Flux and timed seven turns. The eager-to-settled gap ran from 0 to 721 ms, with a median of about 160 ms.

Jev's median response time in the same build was 140 ms over twelve calls, so on some turns the eager call finished before the turn settled and the answer was already waiting at EndOfTurn. On turns where the gap was near zero there was no head start, but the settled pass still reused the eager answer, so the speculation cost nothing extra. You only pay for a discarded call when the settled transcript differs from the eager one. All seven eager transcripts matched their settled transcripts exactly, which is why reuse works, but that's a small sample, so keep the transcript comparison in your code.
In a recorded 50-second session with office chatter playing from a phone across the room, the answer was already waiting on all 11 turns. The win panel counts 46 passes: 35 billed Jev calls, plus 11 settled passes that reused a held answer and made no call. The whole session cost 1.11 cents, Flux included.

Why not act on the eager answer?
The words that change a sentence's meaning often come last, and the eager event fires before they arrive.

Jev Runner uses the none option as a relevance gate and rejects a turn when the probability on none is above 0.5. "We should go left on the pricing page" scores 0.99 on none, so the maze ignores it, but truncate it to "we should go left" and the score drops to 0.22, which would move the marker.
| Full sentence | P(none) | Eager-point truncation | P(none) |
|---|---|---|---|
| we should go left on the pricing page | 0.99 | we should go left | 0.22 |
| lets go right to the point | 0.93 | lets go right | 0.05 |
| push the deadline back three days | 0.99 | push the deadline back | 0.99 |
"Push the deadline back" scored 0.99 either way, because "deadline" arrives before the eager point. Truncation only fools the gate when the giveaway words come last.
I also tested commit rules against about 1,300 logged decisions. The best rule, "commit an eager answer when every move has an explicit count," matched the settled answer on 48% of turns, and tightening the confidence floor didn't move that number. So speculate on eager and commit on settled, because the speculation is cheap and the wrong commit isn't.

Gate on the reject option
When a Choice question has an explicit reject option, gate on that option's probability. Across 11 test utterances, maze commands scored 0.01 to 0.15 on none, while room speech scored between 0.98 and a flat 1.00. Nothing landed in between, which gave me a clear gap to set the threshold in.
Confidence on the same maze commands ranged from 0.82 to 0.98, and it drops for reasons unrelated to relevance. "Right right" sits at 0.82 because the phrasing is odd, while its P(none) stays at 0.15. In a separate run, the gate rejected 12 of 12 room phrases while bare commands like "go left" still moved the marker.

Force the turn in a room that never goes quiet
Flux closes a turn when it hears trailing silence, and a loud room may never give it one. For that case, Flux accepts a ForceEndTurn control message on the same socket:
flux.sendForceEndTurn({ type: 'ForceEndTurn' })In Jev Runner, the period key sends it. In testing, five forced turns each closed 100 to 200 ms after the message, and if no turn is open, Flux ignores the message and returns a FORCE_END_TURN_NO_ACTIVE_TURN warning. Setting eot_threshold to 1.0 turns off natural detection and lets the player own every boundary, which Deepgram's Bring Your Own Turn Detection guide walks through. Leaving it at 0.7 keeps natural detection with the key as an override, which is a happy medium.
The transcript still comes back containing everyone in the room, and Jev decides which words were meant for the maze and which weren't.
What does it cost to run Flux and Jev together?
The usual objection to speculation is the cost in extra LLM calls. Deepgram's own Eager End of Turn guide puts it at "50-70% more LLM calls." This pattern made about three times as many, because Flux revises the eager transcript several times in a turn and each revision asks Jev again. That multiplier is exactly why the model you speculate with has to be cheap.
Here is the measured session from above, at about 13 turns a minute: 11 turns, 35 billed Jev calls, and 11 reused answers. Jev cost 0.56 cents and Flux cost 0.55 cents, at $0.0065 per minute on Deepgram's pricing page, billed on the clock. That works out to about 80 cents an hour, split roughly evenly between the two.
Limits I measured
- Sequence length. Jev reliably pulled three or four moves out of one utterance. Past that, confidence on later moves fell off, and adding more question slots did not help.
- Crosstalk and counts. With other people talking over a counted command, 4 of 6 test turns lost the count.
- Homophones. Flux sometimes transcribes "right" as "write." Naming the homophone in Jev's criteria made the gate worse, so I removed it.
- Browser noise suppression. It treats background speech as noise and filters it out before Flux hears it. For a relevance gate, turn it off.
Start building with Flux eager end of turn
Read the Flux end-of-turn configuration docs for eager_eot_threshold, eot_threshold, and ForceEndTurn. Then move your model call from EndOfTurn to EagerEndOfTurn, hold the answer, and commit on EndOfTurn. Start from the handler above and swap the maze questions for your own. If you don't have a Deepgram API key yet, sign up at console.deepgram.com/signup.
Run it
The whole pattern is in one runnable file: it streams a 16 kHz mono WAV file through Flux in real time and prints Jev's verdict for every turn.
- Download the gist's
flux-eager-jev.tsandpackage.json, then runnpm install. - Set
DEEPGRAM_API_KEYandTYPESAFE_API_KEYin your environment. - Run
node flux-eager-jev.ts room.wavon Node 22.6 or later.