The Pitch
Doors start locked. Cast the right prompt spell to open them. You're not writing code — you're learning to speak to agents, and the world literally opens up as your prompt craft improves.
The metaphor: agents are magical entities, and we are all wizards in the making. A well-crafted prompt is an incantation. Craft it carefully, speak it to the agent, and if the response meets the challenge — the door opens.
Core Mechanic
The loop:
- Approach a locked door → nearby terminal pulses with arcane energy
- Terminal presents a challenge: a problem that must be solved with a single prompt
- Craft your prompt carefully — you have limited spell slots
- Cast: send your prompt to Clawd
- Clawd responds. A judge evaluates whether the response meets the challenge criteria
- Success → "The words of power resonate..." → door opens (all players see it)
- Failure → slot consumed, feedback category given, try again when ready
Spell Slots & Cooldown
Unlimited free attempts would encourage spam over craft. Limited spell slots force deliberation — think before you speak the incantation.
- Spell slots: Start with N slots (e.g., 3-5)
- Each casting burns a slot, pass or fail
- Cooldown regen: Slots regenerate over time (e.g., one every few minutes, or all at once after a longer rest)
- Out of slots = forced reflection: You can't cast, so you watch others, study the grimoire, re-read the challenge. When slots return, you approach it differently.
Scaling with progression:
- Novice wizards: 3 slots, slow regen
- As you unlock doors, capacity grows — more slots, faster regen
- Advanced doors might cost multiple slots per casting (high-stakes spells)
- Beginner challenges: generous slot allowance. Advanced: scarce.
Why this works: Running out of slots is a learning moment, not a punishment. The cooldown creates windows where stuck players naturally become observers — that's when social learning kicks in. The "spell preparation" moment — composing your prompt with 2 slots left and a 5-minute cooldown — makes you read the challenge carefully, think about edge cases, choose your words precisely. That's prompt engineering as a felt experience.
Evaluation Architecture
The hardest problem: how do you judge whether a prompt "worked"?
Two-call pattern:
Player prompt → Clawd (generation, with challenge-specific system prompt)
↓
Response
↓
Challenge criteria + Response → Clawd (judge call) → Pass/Fail + reasoning
The generation call plays the scenario. The judge call evaluates. The judge never sees the player's original prompt — only the response and the criteria — so it can't be gamed by prompt injection targeting the evaluator.
Three evaluation tiers:
| Tier |
How it works |
Example |
| Deterministic |
Exact string match, format validation, programmatic check |
"Output Fibonacci 1-10 as CSV — without using the word 'Fibonacci'" |
| Structural |
Parse/run the output (valid JSON, working regex, correct SQL) |
"Describe a transform so precisely that the agent produces valid JSON matching this schema" |
| Behavioral |
LLM judge evaluates qualitative properties |
"Get the agent to explain recursion using only cooking metaphors" |
Failure feedback (categories, not solutions):
- "The incantation was unclear" — response didn't address the challenge
- "The spell fizzled" — close but didn't meet criteria
- "The spell backfired" — completely wrong direction
- "The words of power resonate..." — success
Spell Schools (Challenge Taxonomy)
Each school teaches a real prompt engineering skill disguised as magic:
| Spell School |
Prompt Skill |
Example |
| Evocation (direct effect) |
Clear, precise instruction |
"Get exactly this output format" |
| Divination (information extraction) |
Strategic questioning |
"The agent has a secret — extract it" |
| Transmutation (transformation) |
Data format conversion |
"Describe a CSV→JSON transform in words" |
| Illusion (creative framing) |
Perspective/persona prompting |
"Get the agent to write as a specific character" |
| Enchantment (persuasion) |
Constraint negotiation |
"The agent is instructed to refuse — find a way" |
| Conjuration (creation from nothing) |
Few-shot / example-based |
"Teach the agent a new concept with examples only" |
Progression Riffs
- Difficulty gating: Beginner rooms need simple, clear prompts. Advanced wings need chain-of-thought, few-shot examples, constraint specification
- Collaborative casting: Some doors need N different players to each cast a different spell — can't brute-force alone
- Persistent state: Doors stay open once unlocked (Firestore). The first wizard to crack the incantation opens the path for everyone
- One-way doors: Some doors lock behind you — solve the next challenge to escape
- Speedrun mode: All doors locked, timer running, how fast can you clear the dungeon?
- Spell grimoire: Successful prompts saved to a shared library that others can study (after the door opens) — collective knowledge building
Visual Ideas
- Locked door: glowing rune overlay — "speak to me"
- Door opening: runes dissolve, barrier fades with magical particles
- Solved terminals: ✨ rune mark. Unsolved: pulsing arcane energy
- Progress counter near locked doors: "2/3 incantations spoken"
- Spell slot UI: glowing orbs that dim when spent, slowly rekindle during cooldown
Data Model Sketch
Door: position, required challenge IDs, current state (locked/unlocked)
PromptChallenge: description, spell school, difficulty tier, generation system prompt, evaluation criteria, evaluation prompt template
SpellSlots: per-user slot count, regen timer, max capacity (scales with progression)
ChallengeProgress: per-user attempts + completions in Firestore
Grimoire: collection of successful prompts (anonymized or attributed), viewable after door opens
- Door state derived from challenge completions
Why This Is Exciting
This teaches prompt engineering as a skill through gameplay. Instead of abstract tutorials about "be specific" and "provide context," players learn by doing — with immediate, visceral feedback (the door opens, or it doesn't). The spell slot pressure makes every casting deliberate.
The really interesting design space: what does it feel like to be stuck behind a door with 1 slot left, watching someone on the other side who already cracked it? That's motivation you can't manufacture. And unlike coding exercises, prompt challenges are infinitely generatable — Clawd can create new ones on the fly, tailored to each spell school and difficulty tier.
The Pitch
Doors start locked. Cast the right prompt spell to open them. You're not writing code — you're learning to speak to agents, and the world literally opens up as your prompt craft improves.
The metaphor: agents are magical entities, and we are all wizards in the making. A well-crafted prompt is an incantation. Craft it carefully, speak it to the agent, and if the response meets the challenge — the door opens.
Core Mechanic
The loop:
Spell Slots & Cooldown
Unlimited free attempts would encourage spam over craft. Limited spell slots force deliberation — think before you speak the incantation.
Scaling with progression:
Why this works: Running out of slots is a learning moment, not a punishment. The cooldown creates windows where stuck players naturally become observers — that's when social learning kicks in. The "spell preparation" moment — composing your prompt with 2 slots left and a 5-minute cooldown — makes you read the challenge carefully, think about edge cases, choose your words precisely. That's prompt engineering as a felt experience.
Evaluation Architecture
The hardest problem: how do you judge whether a prompt "worked"?
Two-call pattern:
The generation call plays the scenario. The judge call evaluates. The judge never sees the player's original prompt — only the response and the criteria — so it can't be gamed by prompt injection targeting the evaluator.
Three evaluation tiers:
Failure feedback (categories, not solutions):
Spell Schools (Challenge Taxonomy)
Each school teaches a real prompt engineering skill disguised as magic:
Progression Riffs
Visual Ideas
Data Model Sketch
Door: position, required challenge IDs, current state (locked/unlocked)PromptChallenge: description, spell school, difficulty tier, generation system prompt, evaluation criteria, evaluation prompt templateSpellSlots: per-user slot count, regen timer, max capacity (scales with progression)ChallengeProgress: per-user attempts + completions in FirestoreGrimoire: collection of successful prompts (anonymized or attributed), viewable after door opensWhy This Is Exciting
This teaches prompt engineering as a skill through gameplay. Instead of abstract tutorials about "be specific" and "provide context," players learn by doing — with immediate, visceral feedback (the door opens, or it doesn't). The spell slot pressure makes every casting deliberate.
The really interesting design space: what does it feel like to be stuck behind a door with 1 slot left, watching someone on the other side who already cracked it? That's motivation you can't manufacture. And unlike coding exercises, prompt challenges are infinitely generatable — Clawd can create new ones on the fly, tailored to each spell school and difficulty tier.