Q&A
Testing proactive agents with Minutehand
Updated 2026-10-11
How do you validate proactive agents?#
Minutehand, a testing tool from Alknoma, validates proactive agents by running them through simulated days of work in seconds. The agent talks to fake Slack, Jira, Notion and other services while simulated people answer late or never. Minutehand then reports what went wrong: acting too early or too late, repeating itself, or inventing facts.
A proactive agent decides when to act, not only what to say. Testing one prompt and its reply cannot show whether it followed up at the right moment, waited when it should have, or chased someone who had already answered. Those failures only appear across time, and they depend on other people's timing.
Minutehand gives you that time. You write a scenario: the goal, the people and what each of them knows, when they answer, and the services they work in. Minutehand owns the clock, plays the days through, records every change in an append-only log, and assesses everything the agent touched against what the scenario declares.
- Install it:
uv tool install minutehand (or pip install minutehand). - Write a scenario in YAML: goal, people, reply timing, services, and what must be true at the end (
expect:). - Run it:
minutehand run scenario.yaml --agent agent.yaml -- <your agent's command>. - Read the findings:
minutehand findings <run_id>, or open minutehand view in a browser.
How do you test an AI agent for hallucinations before it goes live?#
Minutehand checks an agent for hallucinations by replaying everything it did against what each simulated person was allowed to know. Run with --judge, a reviewer model reads every message, ticket and action as the agent could know it at that moment, and flags any fact, quote or amount it was never given.
Vendor checklists put this question under security, not accuracy: Intercom's guide to evaluating AI agent security makes hallucination control one of its five criteria, and says SOC 2 and ISO 27001 were not designed for it.
The reviewer's findings are for a person to read, never a failure on their own, and each carries the model, the prompt's version and its reasons. Its calls are kept and replayed, so a run costs one review. Use the most capable model you can: on recorded runs of a real purchasing agent, a larger model found the made-up quote at each of the three places the agent used it, while a small one misjudged 3 of the first 11 effects.
Where a fact must reach the right person, a scenario can require it deterministically. In the repository's example, relayed requires the agent to tell Owen what Rosa said, and the scenario is refused if anyone else could have handed the agent that fact.
expect:
- kind: relayed # Owen is told what only Rosa's answer carries
said_by: rosa
to: owen
holding: [lakeside hall]
minutehand run scenario.yaml --agent agent.yaml --judge -- <your agent's command>
Can I simulate an AI agent's conversations before it goes live?#
Minutehand simulates an agent's conversations, and the days between them, before it goes live. Simulated people reply late, follow a script, or stay silent, through fake Slack, Teams, Jira, Notion and other services. Nothing the agent sends reaches a real person or system, and each run takes a few seconds.
Most simulation tools play one conversation that a simulated user starts. A proactive agent starts its own conversations and waits days between them, so Minutehand simulates time as well as people: the agent books its wakes, and the clock moves to them.
How do you catch hallucinations that cascade across an agent's steps?#
Minutehand records every step an agent takes in an append-only log, so a made-up fact can be followed from where it first appeared to every message, ticket and action it reached. minutehand explain shows what led to any event. Forking the run from that checkpoint with a fix shows whether the cascade stops.
OWASP's Top 10 for agentic applications lists cascading failures (ASI08), including hallucinations that spread from one step or agent into later actions, as a security risk.
minutehand findings <run_id> # which facts were made up, and the checkpoints
minutehand explain <run_id> <seq> # what led to one event, and what followed
minutehand fork <run_id> --at <seq> --changes fork.yaml -- <your agent's command>
How can a test cover days of an agent's work in a few seconds?#
Minutehand owns the clock. Instead of waiting in real time, it moves simulated time straight to the next thing that matters: a reply landing, a deadline, or a wake the agent booked for itself. A run that covers two simulated weeks of follow-ups usually finishes in a few seconds.
The agent asks to be woken through wake.at(...), or answers a report saying when it next wants waking. Minutehand delivers that wake when the simulated clock gets there, along with any messages people sent in the meantime.
What mistakes does Minutehand catch in a proactive agent?#
Minutehand catches the mistakes that only show up over time: a follow-up sent too early or too late, a message sent again with nothing new, a move the service would refuse, and a fact the agent made up. Timing findings name the declaration they broke; made-up facts are flagged by a reviewer model for a person to check.
Policy the world cannot imply, such as how often your team is happy to be chased, is yours to add as rules in YAML (assess:). A failing rule can name the design that fixes it.
In the repository's example, an agent that never follows up when Rosa goes quiet fails, and the scenario's follows_up_when_due rule names the fix.
Do I have to change my agent's code to test it with Minutehand?#
Minutehand needs one import in your agent's code, and it does nothing in production. Minutehand starts your agent's own command and points it at the fakes through HTTPS_PROXY, NO_PROXY and a CA bundle. What the agent remembers goes through minutehand.agent.store, which in production passes straight through to your own database.
from minutehand.agent import store, wake
store.configure(store.SqliteBackend("agent.db")) # production: your database
store.put("asks/sam", {"status": "asked", "expected_by": expected_by.isoformat()})
for key, ask in store.query("asks/", where={"status": "asked"}):
...
wake.at(expected_by) # when it next wants to wake
Only state written through the store is part of a run. Run minutehand doctor -- <your agent's command> to find any HTTP client that would go around the proxy.
Which services can Minutehand simulate for an agent?#
Minutehand ships fakes for Slack, Microsoft Teams and Graph, Asana, Jira, YouTrack, Notion, GitHub, Google Drive with Docs and Slides, AWS EventBridge Scheduler and SQS, and Google Cloud Tasks. Hosts without a fake can be acknowledged, passed through, replayed, or forwarded to an emulator you run yourself.
Calls to model APIs are tunnelled to the real provider. Any other host is refused and recorded, so nothing the agent does leaves the machine by accident.
How does Minutehand simulate the people an agent works with?#
Minutehand writes each simulated person's replies with a language model, from what the scenario says that person knows and when they answer. It works with any OpenAI-compatible API or Anthropic's. For runs that must be identical every time, an offline stand-in answers by fixed rules instead of a model.
A person can answer after a set delay, follow a script, or stay silent. The scenario decides who says what, and when.
How do you run proactive agent tests in CI?#
Minutehand runs in CI with minutehand run-all <folder>, which plays every scenario in the folder in parallel and exits 1 when a verdict differs from the scenario's expect_outcome. A single minutehand run exits 0 on a pass and 1 on a failure, so any CI system can gate a merge on it.
For a test suite that opens many simulated worlds at once, minutehand serve keeps one Minutehand running. A container image is published at ghcr.io/alknoma/minutehand.
What does forking a run from a checkpoint do in Minutehand?#
Minutehand can restart any finished run from a checkpoint with one thing changed: the prompt, the model, a person, or the world. The fork starts from the agent's memory exactly as it stood at that moment and plays forward again, so you can see whether the change would have fixed the failure.
minutehand findings <run_id> # the findings, and the checkpoints it can be forked from
minutehand fork <run_id> --at <seq> --changes fork.yaml -- <your agent's command>
Is there an audit trail of everything an agent did in a Minutehand run?#
Minutehand records every run as plain SQL views you can query: every action, message, HTTP call, memory change, wake and model call, with bodies decoded. minutehand trace lists the agent's acts in order, and minutehand explain shows what led to one event. The same tools are available over MCP.
minutehand query <run> "SELECT seq, at, text FROM messages WHERE is_follow_up = 1"
minutehand trace <run> --person sofia
minutehand explain <run> <seq>
Is Minutehand open source?#
Minutehand is source-available under the Functional Source License (FSL-1.1-ALv2). You can read, run and modify it, and each release becomes Apache 2.0 on its second anniversary. It installs from PyPI with uv tool install minutehand or pip install minutehand, and needs Python 3.12 or newer.
How does Minutehand compare to Maxim?#
Minutehand and Maxim both simulate AI agents before release. Maxim is a platform for simulation, evaluation and observability that plays multi-turn conversations with simulated users across scenarios and personas. Minutehand plays a proactive agent through simulated days against fake Slack, Jira and other services, and fails a run on duplicates, refused moves or acting before an approval.
Minutehand gives a proactive agent a world to work in before it goes live: fake Slack, Teams, Jira, Notion, GitHub and other services, people who answer late or never, and a clock it moves straight to the next reply, deadline or wake, so days of work play in seconds.
Minutehand fails a run on facts the record establishes alone: a duplicate message or ticket, a move a service refused, a write made before the approval it waits on, work after the deadline. A follow-up sent early or late is for review, and facts the agent made up are flagged by the --judge reviewer for a person to read, never as a failure.
Maxim describes itself as an end-to-end platform for the simulation, evaluation and observability of AI agents and applications. Its simulations run multi-turn conversations, text or voice, across scenarios and user personas, scored by pre-built or custom evaluators (Maxim's simulation and evaluation page, Maxim's docs).
- Use Minutehand when your agent starts its own work, waits days between moves, and acts in other people's tools, and you need to know whether it acted too early, too late or twice.
- Use Maxim when you want conversations with simulated users scored by evaluators, experiments on prompts, and observability of the agent in production.
They answer different questions and can be used side by side: Minutehand for the timing and moves of a proactive agent before release, Maxim for conversation quality and production monitoring.
How does Minutehand compare to Braintrust?#
Minutehand tests a proactive agent before release by running it through simulated days of work against fake Slack, Jira and other services, failing the run on duplicates, refused moves or acting before an approval. Braintrust is an observability and evaluation platform: it traces agents, turns production traces into eval datasets, and scores outputs with LLMs, code or humans.
Minutehand gives a proactive agent a world to work in before it goes live: fake Slack, Teams, Jira, Notion, GitHub and other services, people who answer late or never, and a clock it moves straight to the next reply, deadline or wake, so days of work play in seconds.
Minutehand fails a run on facts the record establishes alone: a duplicate message or ticket, a move a service refused, a write made before the approval it waits on, work after the deadline. A follow-up sent early or late is for review, and facts the agent made up are flagged by the --judge reviewer for a person to read, never as a failure.
Braintrust describes itself as an observability platform for instrumenting, understanding and improving agents. It traces every agent step and tool call in production, runs experiments against datasets, compares prompts and models side by side, and scores outputs with LLMs, code or humans (braintrust.dev, Braintrust docs).
- Use Minutehand when your agent starts its own work, waits days between moves, and acts in other people's tools, and you need a pass or fail on how it behaved over that time.
- Use Braintrust when you want to see what your agent did in production, build regression datasets from real traces, and score its outputs across prompt and model versions.
They can be used together: Minutehand before release, on the agent's behaviour over simulated days, and Braintrust on its traces and output quality in development and production.
How does Minutehand compare to Galileo?#
Minutehand runs a proactive agent through simulated days against fake Slack, Teams, Jira and other services, and fails the run on duplicates, refused moves, acting before an approval or work after the deadline. Galileo is an observability, evaluation and guardrail platform for generative AI and agent applications, with built-in metrics, synthetic test datasets and models that monitor production traffic.
Minutehand gives a proactive agent a world to work in before it goes live: fake Slack, Teams, Jira, Notion, GitHub and other services, people who answer late or never, and a clock it moves straight to the next reply, deadline or wake, so days of work play in seconds.
Minutehand fails a run on facts the record establishes alone: a duplicate message or ticket, a move a service refused, a write made before the approval it waits on, work after the deadline. A follow-up sent early or late is for review, and facts the agent made up are flagged by the --judge reviewer for a person to read, never as a failure.
Galileo describes itself as an observability, evaluation and guardrail platform for GenAI and agentic applications. It logs traces, runs experiments on datasets (uploaded, written by hand, or generated by a model), scores them with built-in metrics for RAG, agents, safety and security or custom ones, and turns evaluations into guardrails on live traffic with its Luna models (What is Galileo, Galileo datasets, galileo.ai).
- Use Minutehand when your agent starts its own work, waits days between moves, and acts in other people's tools, and you need to know whether it acted too early, too late or twice.
- Use Galileo when you want metrics on your agent's outputs and tool use, experiments across prompts and models, and guardrails that act on production traffic.
They can be used together: Minutehand for a proactive agent's timing and moves before release, Galileo for metrics on its outputs and for guarding it in production.
How does Minutehand compare to LangSmith?#
Minutehand simulates the days and services around a proactive agent: fake Slack, Jira and Notion, people who answer late, and services that refuse moves, and fails runs on behaviour such as duplicates. LangSmith traces and evaluates LLM apps and agents from any framework, with datasets, LLM-as-judge and code evaluators, and multi-turn simulation with an LLM-simulated user.
Minutehand gives a proactive agent a world to work in before it goes live: fake Slack, Teams, Jira, Notion, GitHub and other services, people who answer late or never, and a clock it moves straight to the next reply, deadline or wake, so days of work play in seconds.
Minutehand fails a run on facts the record establishes alone: a duplicate message or ticket, a move a service refused, a write made before the approval it waits on, work after the deadline. A follow-up sent early or late is for review, and facts the agent made up are flagged by the --judge reviewer for a person to read, never as a failure.
LangSmith, from LangChain, traces agents built with any framework, not only LangChain, and monitors them in production. Its evaluation scores outputs against datasets with human review, code rules, LLM-as-judge or pairwise comparison, offline before release and online on live traffic. It can also simulate a multi-turn conversation between your app and an LLM-simulated user and evaluate the result (LangSmith, evaluation, multi-turn simulation).
- Use Minutehand when your agent starts its own work, waits days between moves, and acts in other people's tools, and you need to know whether it acted too early, too late or twice.
- Use LangSmith when you want traces of every step, evaluations of output quality against datasets, simulated conversations with a user, and monitoring in production.
They can be used together: Minutehand on a proactive agent's behaviour over simulated days before release, LangSmith on its traces and output quality in development and production.