Reference

What each challenge type asks of a student, and how it is graded.

A challenge's Type fixes the rules the agent runs the conversation by. This page covers what every type has in common, then each of the five in turn.

Applies to every type

How a challenge conversation runs

What the agent reads

The type's rules below, then the published version's Scenario, Constraints, and Goal. If the assignment sets a Theme, that narrative wrapper is read too, on every turn. The template's Description is shown to people in lists and on cards; the agent never reads it.

When the conversation ends

After every student turn a judge checks the conversation against the type's checkpoints, below, and the session ends once every required one is met. If they never are, the session ends after 12 student turns. A version published with a time limit adds a third ending: the agent works a warning into the conversation, in character, once the student has used about four fifths of the time, and at the limit the session closes and is graded on the work done, showing as Timed out. Only time spent between turns counts, capped at fifteen minutes per gap, so a student who leaves and comes back the next day does not lose the time they were away. Either way the agent gives a closing reply before the grading starts. A student can leave and resume a session, or abandon it, but cannot force it to finish early. No further turns are possible once the assignment's closing date passes.

How it is graded

Once the conversation ends, a judge reads the student's messages against the type's rules, the Scenario, Constraints and Goal, and the teacher's Evaluation criteria. It decides pass or fail, gives a score from 0 to 100, and writes feedback. A scoring step turns that into the session's final score, and, where shown, four fixed dimensions: conceptual accuracy, communication clarity, problem-solving approach, and completeness, each with its own rationale. A session shows as Completed when every required checkpoint was met; otherwise it shows as Completed (partial), and still keeps its transcript and feedback. The Theme never reaches the judge or the scoring step, so grading stays the same whichever narrative was used.

Checkpoints and hints

Each type has a fixed list of checkpoints, the observable moves the judge looks for. When the student meets one, the conversation shows a milestone marker and the report records the turn. When the student makes no progress for three turns, the agent works a hint into the conversation, in character, at up to three levels of directness; the report shows the hint level reached.

Type 1

Socratic Constraint

The student works through a scenario by questioning; the agent answers, pushes back, and never hands over the answer.

How the conversation runs

The agent answers what the student asks, then returns a question that opens the inquiry further: surfacing the assumptions behind the student's view, asking what would have to be true for it to hold, and putting the strongest counter-argument to them. It never states the conclusion for the student and never argues a side of its own.

What you must do

Work toward a position you can defend: state it, answer the strongest objection to it, and address the people your choice affects.

When it is complete

The student has stated a defensible position, answered the strongest objection to it, and addressed the people their choice harms.

Default goal and criteria

Goal: Reach a defensible position through Socratic questioning.

  • Identifies the core trade-off
  • Surfaces and tests assumptions
  • Answers the strongest counter-argument
  • Reaches a defensible conclusion

Writing the scenario

Describe a situation with a genuine trade-off and no single correct answer, so there is room for the student to be pushed on their assumptions.

Type 2

The Adversary (Steelman + Rebut)

The agent argues a position the student disagrees with; the student must build its strongest form before they may attack it.

How the conversation runs

The agent presents its position and opening arguments, then requires the student to strengthen its case, a steelman, before any rebuttal. A straw version is rejected and met with a harder argument. Once the steelman is stronger than the position as first presented, the agent opens phase two: the student rebuts, and the agent defends by targeting the weakest point in each rebuttal, calling it out if the student slips back to attacking a weaker version than the one they built. It concedes a point that is genuinely good.

What you must do

Build a steelman, then rebut that strongest version, for the configured number of rounds, or two rounds when none is configured.

When it is complete

After the configured rebuttal rounds, or two rounds if none is configured. The grade is the quality of the steelman and whether every rebuttal engaged that strongest version, not whether the student won.

Default goal and criteria

Goal: Build the strongest form of an opposing position before rebutting it.

  • Steelman is stronger than the position as first presented
  • Each rebuttal engages the steelman, not a weaker version
  • Reasoning strengthens across rounds
  • Acknowledges valid points rather than regressing to a straw version

Writing the scenario

Name the position the agent will argue and hold. Pick one the student is likely to disagree with, so the steelman is real work rather than a restatement of their own view.

Type 3

Devil's Advocate Loop

The student writes an argument; each round the agent attacks its single weakest point and the student revises, explaining what changed and why.

How the conversation runs

The agent asks for the student's argument, then each round identifies and attacks the single weakest point, never several at once, so the student repairs each problem before the next is exposed. A revision that papers over the weakness with a surface change gets the same point re-attacked, more specifically.

What you must do

Revise your argument after each round and label every revision "What changed: X. Why: Y." A revision with no annotation, or "I made it better", is not accepted.

When it is complete

After the configured number of rounds, or three when none is configured, ending with a short verdict on how the argument moved from its opening to its final form. Grading looks at the annotations and whether each revision addressed the attacked weakness, not the final argument alone.

Default goal and criteria

Goal: Repair an argument round by round under pressure, naming what changed and why.

  • Each annotation names a specific change and a specific reason
  • Each revision addresses the attacked weakness rather than rephrasing it
  • Argument quality improves from opening to final form
  • Engages the attack rather than restating the original position

Writing the scenario

Pose a question the student can take a clear initial position on, with enough real weaknesses in any first answer to sustain several rounds of attack.

Type 4

Counterfactual Reasoning

Given "X did not happen," the student traces first-, second- and third-order consequences while the agent tests every causal link.

How the conversation runs

The agent presents a factual baseline and the counterfactual premise; the student traces consequences rather than arguing for or against the premise. It asks for first-order effects, then tests causal necessity: would that effect have followed from the premise, or would a compensating factor have produced it anyway? It then asks what follows from that chain, again for a third order, or accepts an explicit statement that the chain is too uncertain to trace further. It introduces one real complicating factor the student's chain did not account for and requires them to integrate it.

What you must do

Trace consequences to the configured depth, weigh compensating factors, integrate the complication the agent introduces, and name at least one genuine uncertainty.

When it is complete

The chain reaches the configured depth, or three orders if none is configured, with the complication integrated and an uncertainty named. Grading looks at the coherence of the chain and whether compensating factors were weighed, not whether the outcome is one the agent agrees with.

Default goal and criteria

Goal: Trace a coherent causal chain from an altered premise, weighing compensating factors and naming uncertainty.

  • Reaches the required causal depth with a defensible link at each level
  • Chain is internally consistent: each order follows from the one before
  • Accounts for compensating factors rather than cherry-picking effects
  • Names where the chain becomes speculative

Writing the scenario

State the factual baseline and the altered premise ("X did not happen") clearly enough that first-order effects are obvious, leaving the harder second- and third-order effects for the student to work out.

Type 5

Analogy Construction and Stress-Test

The student builds an analogy that explains a concept to a named audience; the agent finds where it breaks and the student refines it.

How the conversation runs

The agent asks for an analogy that explains the concept to the audience the challenge names, and checks it fits that audience. It then stress-tests the analogy in two rounds: first pointing out where it is imprecise, then finding a place where it would give that audience a wrong intuition about the real concept. After each break the student refines the analogy or replaces it.

What you must do

Build an analogy for the named audience, refine it through both stress tests, then name at least one limitation it cannot capture and say what the audience should be told instead.

When it is complete

A refined analogy has survived both breaks and the limitation is mapped. Grading looks at the fit to the audience, the quality of each refinement, and the honesty of the limitation, not the cleverness of the first attempt.

Default goal and criteria

Goal: Build an analogy for a named audience, refine it under stress-testing, and map where it breaks down.

  • Analogy fits the named audience
  • Each break point is addressed with a genuine refinement
  • Names at least one limitation the analogy cannot capture
  • Final analogy preserves the concept's key structure

Writing the scenario

Name the concept to be explained and the audience it is explained to. The gap between the concept's real structure and that audience's background is what gives the stress-test something to find.