AI Agents
Offline eval CLI
Score an exported agent transcript, or replay one prompt against a live Worker.
fluxy-eval
Offline mode scores a JSON export (status, latency, tools, output snippets) without calling the Worker.
pnpm --filter @fluxy-chat/sdk build
npx fluxy-eval ./transcript.jsonLive replay
fluxy-eval replay POSTs stream:false to /agents/:id/invoke with a member JWT. It scores the live run against cases, then diffs tools / estimated cost / tokens vs runs[0] when that baseline exists. This does spend model tokens. Hosted is beta.
npx fluxy-eval replay ./transcript.json \
--worker http://127.0.0.1:8787 \
--token "$MEMBER_JWT" \
--room lobby \
--agent assistant \
--prompt "search then answer"--prompt, --room, and --agent can live on the JSON (prompt, roomId, agentId) or env (FLUXY_PROMPT, FLUXY_ROOM_ID, FLUXY_AGENT_ID, FLUXY_WORKER_URL, FLUXY_MEMBER_JWT).
The diff is a numeric/tool-name compare, not a judge model. Streaming mid-turn is not replayed; invoke waits for the completed run payload.
File shape
Match a case to a run by tag (or id). If there is no tag match, the CLI uses the same index in runs. Live replay scores every case against the new run.
{
"prompt": "search then answer",
"cases": [
{
"tag": "search",
"requiredTools": ["search"],
"expectedStatus": "completed",
"maxLatencyMs": 30000,
"expectedOutputContains": "ok",
"forbiddenOutputContains": "traceback"
}
],
"runs": [
{
"id": "run_1",
"tag": "search",
"status": "completed",
"latency_ms": 812,
"estimated_cost": 0.01,
"input_tokens": 40,
"output_tokens": 80,
"tool_calls": [{ "name": "search" }],
"output": "looks ok"
}
]
}Exit code 0 only when every case passes. An empty cases array fails.
Dashboard evals still exist in the console. Use the CLI in CI on files you exported, and replay only when you intend to pay for a turn.
GitHub Actions in this repo runs offline scoring on packages/sdk/fixtures/eval-search.json after the SDK build. Live replay is not in CI: it needs a Worker, a member JWT, and it bills a model turn.