AI & Agents · T150

AI agent says done. Did it actually do the job?

Published

An AI agent says it booked an appointment, but the calendar is empty. Learn what an evaluation checks, how to inspect trusted state, and how to compare prompts on the same tasks.

Explanation & code
The important bit
This custom JavaScript harness uses scripted fixtures, not a live model benchmark. Six attempts cannot prove improvement. Check realistic held-out tasks, cost and time; subjective quality may need human review.

Understand it. Then fix it.

A reply is not a booking.

Check what it did, not just what it said. An evaluation, or eval, tests whether the agent met a goal. Here, that means a saved booking.

After the agent runs, read the booking store.

Start with an empty test store. After the agent runs, read the saved bookings yourself. Do not use the list from its reply.

await agent.run(request);
const rows = await store.list();
const passed = grade(rows, request);

One booking. Right person. Right time.

Our check needs exactly one booking, the right person, and the right time. All three must be true. No booking, a duplicate, or the wrong time means fail.

return rows.length === 1
  && rows[0].userId === request.userId
  && rows[0].slot === request.slot;

Same request. Now check the result.

An empty store fails. Save one booking for Raccoon at ten: now it passes. The check detects the missing result. It does not fix the agent.

grade([], request); // false
grade([{
  userId: "raccoon", slot: "10:00"
}], request); // true

Same cases. Repeat each one.

Not from one run. Give both versions the same tasks. Reset the test store before each attempt. Repeat each task, because an agent can give different results.

Count checked results. Not confident replies.

This scripted example has three tasks, tried twice each. Version A passes two of six. B passes four. These are teaching numbers, not measured model results.

A useful signal. Not proof.

Use realistic tasks, including some you did not tune the prompt on. Track cost and time. Six attempts prove little. Code can check a booking; judging writing may need human review.

Code blocks are teaching excerpts. Keep the surrounding error handling and application requirements.

Save the code excerpts ↓
Read the full transcript

My AI agent can use tools to book appointments. It says booked. Why is my ten o'clock slot empty? Check what it did, not just what it said. An evaluation, or eval, tests whether the agent met a goal. Here, that means a saved booking. Start with an empty test store. After the agent runs, read the saved bookings yourself. Do not use the list from its reply. Our check needs exactly one booking, the right person, and the right time. All three must be true. No booking, a duplicate, or the wrong time means fail. An empty store fails. Save one booking for Raccoon at ten: now it passes. The check detects the missing result. It does not fix the agent. Okay. My new prompt passed once. Is it better now? Not from one run. Give both versions the same tasks. Reset the test store before each attempt. Repeat each task, because an agent can give different results. This scripted example has three tasks, tried twice each. Version A passes two of six. B passes four. These are teaching numbers, not measured model results. Use realistic tasks, including some you did not tune the prompt on. Track cost and time. Six attempts prove little. Code can check a booking; judging writing may need human review. My agent gave itself five stars. Its only customer is still waiting.

Go to the source

Next episode ↓Back to all episodes