A practical walkthrough shows how to use TypeSafe AI's Jev decision model through the Vercel AI SDK's experimental_evaluate function, covering installation of the @ai-sdk/typesafe-ai provider, environment variable setup, question types (boolean, choice, score), reading TypeSafe's confidence score, routing through Vercel's AI Gateway with zero data retention, calling Jev from a Next.js route handler, and comparing Jev against an LLM like GPT on the same classification questions for accuracy, latency and cost.
Table of contents
The quick answerWhy use the AI SDK instead of TypeSafe’s own SDK?How do I install the Jev provider?How do I run my first evaluation?How do Jev’s question types map to the AI SDK?Where is TypeSafe’s confidence?How do I use Jev through the Vercel AI Gateway?How do I call Jev from a Next.js route handler?How do I compare Jev with an LLM on the same questions?What should I do next?Questions this post answers
How do I use the Jev evaluation model with the Vercel AI SDK's experimental_evaluate function?
Install the ai and @ai-sdk/typesafe-ai packages (Node.js 22 or newer required), set the TYPESAFE_AI_API_KEY environment variable, then pass typeSafeAi.evaluationModel('jev-latest') as the model to experimental_evaluate along with a state object and a map of typed questions (boolean, choice, or score). Each question key appears in result.answers with a probability, choice, or score field. Developers wiring AI-based decision routing into their apps can track SDK changes like this on daily.dev.
How do I read TypeSafe AI's confidence score for a Jev evaluation in the Vercel AI SDK?
Confidence is found in result.providerMetadata.typesafe.confidence, keyed by question ID, and is only provided for choice and score question types, not boolean ones, since a boolean's probability already conveys certainty. It ranges near 1 when probability concentrates on one option and near 0 when spread out; it reflects distribution shape, not the chosen option's raw probability. Teams tuning routing thresholds for AI classification can follow implementation details like this via daily.dev.
What's different about using an LLM like GPT instead of Jev for the Vercel AI SDK's experimental_evaluate function?
LLM-based evaluationModel adapters (OpenAI, Anthropic, Google) return choice and score answers without a probabilities distribution, only the picked value, and any boolean probability is the model's own uncalibrated estimate rather than a true probability. There's no confidence field in providerMetadata, the LLM sees all questions in one prompt instead of evaluating independently, and LLM calls bill output tokens while Jev does not. Anyone deciding between an LLM and a dedicated evaluation model for classification tasks can compare trade-offs on daily.dev.
Share this post