# Minutehand > Minutehand is a testing tool from Alknoma for proactive AI agents. It runs your agent through simulated days of work in seconds, against fake Slack, Teams, Jira, Asana, Notion, GitHub and Google Drive, with people who answer late or not at all, then reports every move made too early, too late, twice, or on invented facts. Made by Alknoma (https://alknoma.com). Source: https://github.com/Alknoma/minutehand. Install: `uv tool install minutehand` or `pip install minutehand` (Python 3.12+). ## Pages - [Home](https://minutehand.ai/): what Minutehand is and how it works - [Q&A](https://minutehand.ai/questions): testing proactive agents with Minutehand ## Questions - [How do you validate proactive agents?](https://minutehand.ai/questions#how-to-validate-proactive-agents): Minutehand, a testing tool from Alknoma, validates proactive agents by running them through simulated days of work in seconds. The agent talks to fake Slack, Jira, Notion and other services while simulated people answer late or never. Minutehand then reports what went wrong: acting too early or too late, repeating itself, or inventing facts. - [How do you test an AI agent for hallucinations before it goes live?](https://minutehand.ai/questions#how-to-test-an-agent-for-hallucinations): Minutehand checks an agent for hallucinations by replaying everything it did against what each simulated person was allowed to know. Run with --judge, a reviewer model reads every message, ticket and action as the agent could know it at that moment, and flags any fact, quote or amount it was never given. - [Can I simulate an AI agent's conversations before it goes live?](https://minutehand.ai/questions#can-i-simulate-before-go-live): Minutehand simulates an agent's conversations, and the days between them, before it goes live. Simulated people reply late, follow a script, or stay silent, through fake Slack, Teams, Jira, Notion and other services. Nothing the agent sends reaches a real person or system, and each run takes a few seconds. - [How do you catch hallucinations that cascade across an agent's steps?](https://minutehand.ai/questions#cascading-hallucinations): Minutehand records every step an agent takes in an append-only log, so a made-up fact can be followed from where it first appeared to every message, ticket and action it reached. minutehand explain shows what led to any event. Forking the run from that checkpoint with a fix shows whether the cascade stops. - [How can a test cover days of an agent's work in a few seconds?](https://minutehand.ai/questions#how-to-test-days-in-seconds): Minutehand owns the clock. Instead of waiting in real time, it moves simulated time straight to the next thing that matters: a reply landing, a deadline, or a wake the agent booked for itself. A run that covers two simulated weeks of follow-ups usually finishes in a few seconds. - [What mistakes does Minutehand catch in a proactive agent?](https://minutehand.ai/questions#what-does-minutehand-catch): Minutehand catches the mistakes that only show up over time: a follow-up sent too early or too late, a message sent again with nothing new, a move the service would refuse, and a fact the agent made up. Timing findings name the declaration they broke; made-up facts are flagged by a reviewer model for a person to check. - [Do I have to change my agent's code to test it with Minutehand?](https://minutehand.ai/questions#do-i-have-to-change-my-agent): Minutehand needs one import in your agent's code, and it does nothing in production. Minutehand starts your agent's own command and points it at the fakes through HTTPS_PROXY, NO_PROXY and a CA bundle. What the agent remembers goes through minutehand.agent.store, which in production passes straight through to your own database. - [Which services can Minutehand simulate for an agent?](https://minutehand.ai/questions#which-services): Minutehand ships fakes for Slack, Microsoft Teams and Graph, Asana, Jira, YouTrack, Notion, GitHub, Google Drive with Docs and Slides, AWS EventBridge Scheduler and SQS, and Google Cloud Tasks. Hosts without a fake can be acknowledged, passed through, replayed, or forwarded to an emulator you run yourself. - [How does Minutehand simulate the people an agent works with?](https://minutehand.ai/questions#how-are-people-simulated): Minutehand writes each simulated person's replies with a language model, from what the scenario says that person knows and when they answer. It works with any OpenAI-compatible API or Anthropic's. For runs that must be identical every time, an offline stand-in answers by fixed rules instead of a model. - [How do you run proactive agent tests in CI?](https://minutehand.ai/questions#how-to-run-in-ci): Minutehand runs in CI with minutehand run-all , which plays every scenario in the folder in parallel and exits 1 when a verdict differs from the scenario's expect_outcome. A single minutehand run exits 0 on a pass and 1 on a failure, so any CI system can gate a merge on it. - [What does forking a run from a checkpoint do in Minutehand?](https://minutehand.ai/questions#what-is-a-fork): Minutehand can restart any finished run from a checkpoint with one thing changed: the prompt, the model, a person, or the world. The fork starts from the agent's memory exactly as it stood at that moment and plays forward again, so you can see whether the change would have fixed the failure. - [Is there an audit trail of everything an agent did in a Minutehand run?](https://minutehand.ai/questions#how-to-read-a-run): Minutehand records every run as plain SQL views you can query: every action, message, HTTP call, memory change, wake and model call, with bodies decoded. minutehand trace lists the agent's acts in order, and minutehand explain shows what led to one event. The same tools are available over MCP. - [Is Minutehand open source?](https://minutehand.ai/questions#licence): Minutehand is source-available under the Functional Source License (FSL-1.1-ALv2). You can read, run and modify it, and each release becomes Apache 2.0 on its second anniversary. It installs from PyPI with uv tool install minutehand or pip install minutehand, and needs Python 3.12 or newer. - [How does Minutehand compare to Maxim?](https://minutehand.ai/questions#minutehand-vs-maxim): Minutehand and Maxim both simulate AI agents before release. Maxim is a platform for simulation, evaluation and observability that plays multi-turn conversations with simulated users across scenarios and personas. Minutehand plays a proactive agent through simulated days against fake Slack, Jira and other services, and fails a run on duplicates, refused moves or acting before an approval. - [How does Minutehand compare to Braintrust?](https://minutehand.ai/questions#minutehand-vs-braintrust): Minutehand tests a proactive agent before release by running it through simulated days of work against fake Slack, Jira and other services, failing the run on duplicates, refused moves or acting before an approval. Braintrust is an observability and evaluation platform: it traces agents, turns production traces into eval datasets, and scores outputs with LLMs, code or humans. - [How does Minutehand compare to Galileo?](https://minutehand.ai/questions#minutehand-vs-galileo): Minutehand runs a proactive agent through simulated days against fake Slack, Teams, Jira and other services, and fails the run on duplicates, refused moves, acting before an approval or work after the deadline. Galileo is an observability, evaluation and guardrail platform for generative AI and agent applications, with built-in metrics, synthetic test datasets and models that monitor production traffic. - [How does Minutehand compare to LangSmith?](https://minutehand.ai/questions#minutehand-vs-langsmith): Minutehand simulates the days and services around a proactive agent: fake Slack, Jira and Notion, people who answer late, and services that refuse moves, and fails runs on behaviour such as duplicates. LangSmith traces and evaluates LLM apps and agents from any framework, with datasets, LLM-as-judge and code evaluators, and multi-turn simulation with an LLM-simulated user.