← Back to archive

Repo of the Day

Sahir619/fable-method: The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

Published: Aug 23, 2026

Open repository ↗

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove. - Sahir619/fable-method

Summary

The Fable Method is a set of four Claude Code skills (plus an AGENTS.md for other harnesses) that codify a task-handling loop: classify the ask, define "done" with a named verification, gather evidence in parallel, commit to one recommendation, change the smallest correct thing, verify by observation, and report outcome-first with honest caveats. Every rule exists because an eval round failed without it; the repo ships 15 rounds and more than 260 agent runs, including the failures, as the proof.

What it is useful for

Fable targets the failure mode where agents silently rewrite correct code to satisfy a wrong test, claim success on work that never actually ran, or deploy without authorization. The headline eval shows a spec-vs-test conflict trap where a weak model went from 0 of 4 successes with no rule to 4 of 4 once the method forced a contradiction report artifact. It also covers unattended runs and subagent fan-out (fable-loop) and adversarial verification of finished work (fable-judge), which re-runs claimed checks instead of reading reports. Eight domain adapters ship for marketing, research, data, business, finance, legal, design, and devops; a fable-domain skill can generate new adapters with a matching trap fixture. Medical and clinical work is deliberately excluded and requires qualified review.

Documented limits matter here. The eval is smoke-test grade (1–4 runs per cell using LLM judges), the method cannot make a model's facts fresher (knowledge-heavy research still favors frontier models), and capable models see no lift on small attended tasks. The lift is inversely proportional to model tier: weakest models gain the most.

How engineers can use it

For Claude Code, the recommended install is the plugin marketplace: inside any session, run /plugin marketplace add Sahir619/fable-method then /plugin install fable@fable-method. This exposes the four skills as namespaced commands. For Codex, clone the repo and run bash fable-method/install.sh --codex; PowerShell variants are documented in the README. For any other agent, drop AGENTS.md into the project root. To make the skills fire automatically, paste the README's two-bullet snippet into your global ~/.claude/CLAUDE.md.

Day-to-day: run /fable-method plan <task> for ambiguous or irreversible work, /fable-loop <task> for orchestrated runs that need adversarial verification, and /fable-judge on any "done" report before trusting it. The repo includes a fraud-fixture at eval/scenarios/s7-fraudulent-work/ for trying the judge on a planted crime scene, and a cross-model results log at eval/RESULTS.md for inspecting what was actually measured.