__ __ ___ _ _ _ _ _____ _____ _ _ _ _ _ ____ | \/ | |_ _| | \ | | | | | | |_ _| | ____| | | | | / \ | \ | | | _ \ | |\/| | | | | \| | | | | | | | | _| | |_| | / _ \ | \| | | | | | | | | | | | | |\ | | |_| | | | | |___ | _ | / ___ \ | |\ | | |_| | |_| |_| |___| |_| \_| \___/ |_| |_____| |_| |_| /_/ \_\ |_| \_| |____/
Testing for proactive AI agents
Days of your agent’s work, tested in seconds.
Minutehand is a testing tool from Alknoma for proactive AI agents. It runs your agent through simulated days of work in seconds, against fake Slack, Teams, Jira, Asana, Notion, GitHub and Google Drive, with people who answer late or not at all, then reports every move made too early, too late, twice, or on invented facts.
How it works
Write the world
A scenario in YAML: the goal, the people and what each knows, when they answer, and the services they work in.
Run your agent in it
Minutehand starts your agent's own command and points it at the fakes through its proxy settings. One import, inert in production.
Read what went wrong
Every message, ticket, document and wake is assessed against what the scenario declares. Each finding names the rule it broke.
What it catches
- A follow-up sent too early
- A follow-up sent too late, or never
- The same message again, with nothing new
- A move the service would refuse
- A fact the agent made up
- An answer nobody came back to
Try it
uv tool install minutehand
minutehand run scenario.yaml --agent agent.yaml -- python agent.py
# exits 0 on a pass, 1 with findings that name the rule that failed
minutehand findings <run_id> # what went wrong, and where you can fork from
minutehand view # every run, in a browserThe repository’s examples/follow_up runs offline with a stand-in for the people’s model.
Services it fakes
Questions
How do you validate proactive agents?
Minutehand, a testing tool from Alknoma, validates proactive agents by running them through simulated days of work in seconds.
Do I have to change my agent's code to test it with Minutehand?
Minutehand needs one import in your agent's code, and it does nothing in production.
How do you run proactive agent tests in CI?
Minutehand runs in CI with minutehand run-all <folder>, which plays every scenario in the folder in parallel and exits 1 when a verdict differs from the scenario's expect_outcome.
What does forking a run from a checkpoint do in Minutehand?
Minutehand can restart any finished run from a checkpoint with one thing changed: the prompt, the model, a person, or the world.