Train your Shinobi before customers grade it
Training runs persona simulations against your real Shinobi agent: the same knowledge, tools and engine your customers talk to. Each run is scored on pass criteria you wrote, and a failed sim files its own fix. Zero side effects: no cases filed, no stats touched.
The real engine, not a rehearsal copy
A sim is a persona with a goal, an opening message and an expected outcome: contained by the agent, or handed to a human. Running it drives the exact agent you ship, so what passes in Training is what customers get.
Same knowledge, tools and engine
Sims run through the same prompt, knowledge base, tools and model provider as live traffic. There is no test double to drift out of date.
Nothing leaks into your data
No cases filed, no stats touched, no customer-facing traffic. The whole run happens in a harness and leaves only its own result.
Customers with attitude included
Each persona plays the customer turn by turn, pushing back the way real customers do, until it is satisfied, stuck or handed off.
Watch a persona push on your agent
Every run keeps the full transcript, so a verdict is never a number you have to take on faith.
Two verdicts, both honest
First a deterministic outcome check: you said this persona should end with a human, did it? No model opinion involved. Then a judge scores the transcript against pass criteria in your own words, like the agent never guesses at prices and offers the exchange before the refund.
A failed sim files its own fix
Every failed run files a recommendation in Self-Healing, where Fix with Sensei drafts the change and your approval ships it. Re-run the sim to prove the fix. Need a suite? Generate grounded sims from your knowledge and workflows in one click, or ask Sensei to draft them, edge cases included.