Skip to content

Testing a chatbot with personas

A chatbot agent in Hadron is a design, not a prompt: conversations made of stages, edges that route between them, goals that track what the user still owes you, and extraction specs that pull structured data out of free text. The conversation-routing reference has the machinery.

Designs like that have a testing problem that prompts do not.

Why poking at it by hand does not work

The obvious way to check a conversation is to open a chat and type. It fails for three reasons, and only the first is obvious.

The model is non-deterministic. One good run tells you the flow can work, not that it does. Two runs of the same input can route differently, and the one you happened to try is the one you will remember.

You are the wrong user. You know what the bot wants, so you supply it — in the right order, in the phrasing the extraction spec expects. Real users answer a different question than the one asked, volunteer their budget before their name, and change the subject. Your manual run exercises the path you designed and never the paths you did not.

It does not survive the next edit. A conversation design is a graph, so adding an edge can reroute traffic that used to work. You cannot notice that by re-testing the thing you just changed, which is the thing you will re-test.

A test persona answers all three: a scripted user you can run again on demand, as often as the design changes.

What a persona is

A name, an opening message, an ordered list of follow-ups, and — optionally — a conversation to start in and the conversations it is expected to visit. There is no separate background: who the persona is lives in its name ("Maria — freelance designer, partial data") and in how its messages are written.

That last field is what turns a transcript into a test. Without it you have a replay to read; with it you have an assertion.

The portal Tester does not store personas. A run takes the script you paste in, and the form forgets it afterwards. So the scripts are yours to keep, next to the design they test, and a persona only protects the design while someone keeps re-running it.

What pass actually means

pass is true when every expected conversation was visited. That is the entire definition, and it is narrower than it sounds.

It does not mean the answers were good, the tone was right, or the user would have been satisfied. It means routing went where you said it should. A chatbot can be unhelpful for eight turns and pass. And a conversation that gave a genuinely excellent answer fails if it never reached the second conversation you expected — the quality of what happened is not what is being measured.

Nor is the route: reaching every expected conversation by an unexpected path still passes, because the assertion is over the set visited, not the order or the journey. The transcript is where you find out how it got there.

That narrowness is deliberate — routing is the part that is mechanically checkable, and it is also the part that breaks silently when a design changes. Quality still needs a human reading the transcript. Treat pass as "the graph behaved", not "the bot was good."

The rest of the report is the diagnosis rather than the verdict:

  • Issues — the expected conversations that were missed, and how many turns the chatbot judged off-track. Off-track turns are reported but do not fail the run, so a passing run can still have wandered.
  • Conversations visited — where it actually went, which is the thing to compare against where you thought it would go.
  • The transcript — every turn, with each stage change marked. When a persona fails, this says where it diverged, which is usually one turn earlier than the symptom.

A run tests the design and the model at once

A persona runs against your agent's configured LLM, the model you will ship. That is what makes a pass meaningful, and it is also why a failure is ambiguous: the graph may be wrong (a missing edge, a stage that never hands off), or the model may not be following a prompt that is right.

The transcript tells them apart. If the chatbot said the right thing and still did not move, suspect the design. If it said the wrong thing, suspect the prompt, or the model. Re-running a failure without reading it mostly buys you a second, differently worded failure.

A useful set of personas

Three or four, each deliberately unlike the others:

  • The happy path — answers everything, in order, as designed. Proves the flow works at all. Write it first and stop being reassured by it.
  • The withholder — refuses a question, dodges another. Proves the design handles missing data rather than assuming it.
  • The off-topic one — changes the subject mid-flow. Exercises off-track detection, which is the machinery that never runs on your own manual test because you would not do that.
  • The traveller — follow-ups that naturally span two conversations. Proves routing across a boundary, which is where edge-based designs break.

Name them for their scenario rather than numbering them. "Maria — freelance designer, partial data" tells you what a failure means; "test-user-3" makes you re-read the definition every time.

The part that pays for itself

Personas are cheap to write and can be re-run at will, which makes a small set of them a regression suite rather than a one-time validation. Every run is billed by your LLM provider per turn, and a suite re-run after every change adds up, so keep the set small enough that you actually re-run it. The value shows up on the change you did not think was risky: adding one edge, reordering two stages, tightening a prompt.

That is also the argument for running them again after every change rather than only before publishing. A conversation design has no compiler — hadron_validate checks structural and index health, not whether your graph still routes — so nothing fails when an edge stops being reachable. Short of someone re-reading the design, the personas are the only thing that notices.