Why this is required?
A restaurant booking agent reads a message such as "table for four on Friday around eight, and one of us is vegetarian" and has to turn it into a booking it can stand behind. Each step can go wrong, and the costs are concrete: a double-booked table, a group confirmed for a room that cannot seat them, or a guest told a slot is free when it is not.
- Intent: new booking, change, cancellation, or a menu question.
- Requested slot: the date and time, checked against opening hours.
- Availability: whether the party size fits at that time.
- Next action: confirm, offer an alternative, ask one question, or hand off.
- Hand-off: large groups, allergy or accessibility needs, and complaints.
The common approach asks a chat model to answer in text and then parses that text. Typesafe's announcement describes the weakness directly: existing LLM outputs are strings that "need to be parsed + validated", and there is "always some risk that the AI goes off the rails".
Diagram source (Mermaid)
flowchart TD
A[Guest message] --> B[Chat model writes free text]
B --> C[Software parses the text]
C --> D[Fields checked after the fact]
D -->|fields valid| E[Booking acted on]
D -->|malformed| F[Retry or fail]
F --> G[Guest waits or gets a wrong answer]
E --> H[Risk of double booking or wrong party size]Every failure in that loop lands on the guest. For a booking agent, the reliability of each decision matters more than the fluency of the reply.
How you can achieve this?
Four steps turn the booking decision into something software can trust.
- Define the output shape before calling the model, so every answer is a value your code already knows how to handle.
- Ask for calibrated confidence on each decision, so the agent knows when to act and when to ask.
- Map confidence to an action with an explicit policy, and keep the restaurant's rules in code.
- Replay past enquiries with known outcomes before going live.
Typesafe describes Jev as "a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out." The shape below illustrates the decisions a booking agent needs. It is not Typesafe's API, so check the current documentation for the exact request and response format.
type BookingIntent = 'new_booking' | 'change' | 'cancel' | 'question';
type BookingAction = 'confirm' | 'offer_alternative' | 'ask_clarifying_question' | 'hand_off';
interface Decided<T> {
value: T;
confidence: number; // calibrated: 0.9 should be right about 9 times in 10
}
interface BookingDecision {
intent: Decided<BookingIntent>;
action: Decided<BookingAction>;
requestedSlot: Decided<string | null>; // local time, e.g. '2026-10-16T20:00'
}Because the values come from a closed set, a reply such as "sometime next week, maybe" has to resolve to a listed option or to a clarifying question. Your code never parses free text.
Diagram source (Mermaid)
flowchart TD
A[Guest message and booking state] --> B[Typed decision model]
B --> H{Hand-off needed}
H -->|large group, allergy or complaint| I[Hand off to a person]
H -->|no| C{Intent and slot confident}
C -->|no| E[Ask one clarifying question]
C -->|yes| D[Check availability in code]
D -->|slot free| F[Confirm booking]
D -->|slot full| G[Offer nearest alternatives]A starting policy
- Hand off: a large group, an allergy or accessibility need, a complaint, or repeated low confidence.
- Ask one question: intent or time is uncertain. Ask one, not several.
- Offer alternatives: the slot is full, chosen from the live availability list.
- Confirm: intent and slot are confident, and the availability check passes.
Start cautious. A wrong confirmation costs more than an extra question, so set the thresholds to ask more often at first, then loosen them with measured results.
Availability, opening hours, maximum party size and deposit rules should come from the restaurant's data and be checked in code after the model responds. A confirm from the model is only a suggestion until that check passes.
With/Without this concept?
The difference is easiest to see side by side. The table compares the two approaches on the points that matter for a booking agent.
| Aspect | Without typed decisions | With typed decisions |
|---|---|---|
| Output | Free text that software must parse | Values from a predefined set |
| Validation | Checked after the fact, and can fail | Built into the output shape |
| Uncertainty | Often unstated, and models tend to be overconfident | Calibrated confidence on each decision |
| Failure path | Retries, guesses, or a broken booking | Clarify, offer alternatives, or hand off |
| Rules | Mixed into prompts | Kept in code and checked after the decision |
| Trade-off | Full language flexibility | No free-text generation, so guest-facing replies still need a language model |
Diagram source (Mermaid)
flowchart LR
subgraph Without typed decisions
W1[Free text reply] --> W2[Parse and hope]
W2 --> W3[Retry on error]
end
subgraph With typed decisions
T1[Typed decision] --> T2[Shape already valid]
T2 --> T3[Confidence picks the action]
endTypesafe says Jev cannot produce type errors, and describes that as a property of its structured outputs rather than a benchmark result. Without the typed layer, you carry the validation, retry and audit logic yourself. With it, the decision arrives in shape, and your code spends its effort on the restaurant's rules.
The trade-off is real. Jev gives up free-text generation, so a booking agent that writes replies to guests still needs a language model for wording. Use a typed decision for the choice, and free text only where the wording matters. Jev was in early access at announcement, so confirm its current availability and API before committing to a build.
