Advisor and orchestrator protocols for Claude Fable and GPT-6 Astra.
Each plugin is a small fable about using a premium model well, sharing a moral — judgment belongs on the strongest model, tokens belong on the cheapest one that can do the leg. Same economics, opposite directions:
| Plugin | Direction | For sessions running on | One line |
|---|---|---|---|
| fable-advisor | consult up | Codex or Claude Code | Buy a compact Fable verdict via Claude CLI or a native Claude agent |
| fable-orchestrator | delegate down | Fable 5.1 (or any premium model, e.g. Opus 5.5) | Plan big, execute small: push bulk reading to cheap parallel workers, keep only distilled findings at Fable rates |
| astra-advisor | consult up | Codex, or Claude Code with Codex CLI | Buy a compact GPT-6 Astra verdict through the host-appropriate route |
| astra-orchestrator | route deliberately | GPT-6 Astra in Codex | Use Sol 6.1 and Luna 6 readers, selectively use Astra, and keep synthesis with the parent |
Install the Fable plugins in Claude Code:
/plugin marketplace add czlonkowski/fables
/plugin install fable-advisor@fables
/plugin install fable-orchestrator@fables
The Astra pair separates compact consultation from scoped reading delegation, with explicit GPT model selection and separate runtime routes:
| Skill | In Codex | In Claude Code |
|---|---|---|
astra-advisor |
Native gpt-6-astra sub-agent, fresh context, high reasoning by default |
Authenticated codex exec --model gpt-6-astra, read-only shell sandbox, ephemeral session |
astra-orchestrator |
Astra parent with native Sol 6.1 / Luna 6 workers or a justified Astra branch | Not supported; CLI consultation belongs to the advisor |
The Astra advisor model ID was checked against the local Codex catalog and official OpenAI
model documentation on
2026-09-07. CLI examples were checked against codex-cli 0.153.4 --help; account
access must still be available when invoked. A skill does not switch the parent
model, and a Claude Code agent's model field cannot select a GPT model.
The orchestrator now prefers Sol 6.1 for code analysis and interpretation, and Luna 6 for bounded extraction when the host exposes it. Resolve exact model IDs and supported effort levels from the current native tool. Terra 5.6 is a legacy option, not a default. This routing update has not been runtime-tested.
This checkout includes .agents/skills/astra-advisor and
.agents/skills/astra-orchestrator as relative symlinks to the packaged skills.
Open Codex in this repository (fable-advisor/) and invoke:
Use $astra-advisor to review this architecture decision before we implement it.
Use $astra-orchestrator to inventory all API clients and their retry behavior.
Codex supports repository skills and symlinked skill
folders. For another project, copy the
complete desired skill folder from plugins/<name>/skills/<name>/ into that
project's .agents/skills/. Both packages also contain .codex-plugin/plugin.json
for Codex plugin distribution. No user-level installation or settings change is
performed by this repository change.
Native delegation explicitly pins model and reasoning and starts with a compact brief instead of the parent transcript. If native sub-agents are unavailable, report that limitation; do not start nested CLI workers from Codex. See the runtime references linked from each skill for the available-tool contract.
The Claude Code marketplace includes astra-advisor alongside the Fable plugins.
To load the plugin directly from this checkout:
claude --plugin-dir ./plugins/astra-advisorOr install it from the marketplace:
/plugin marketplace add czlonkowski/fables
/plugin install astra-advisor@fables
Invoke /astra-advisor:astra-advisor or explicitly ask to use the Astra advisor.
The skill uses your installed, authenticated Codex CLI, passes the compact brief
through stdin, and reads the final verdict file after successful completion. It
does not define a Claude sub-agent with an unsupported GPT model field.
- Advisor: default one consult, maximum three executed interactions per task; brief under 1,200 words, up to five file pointers, answer under 300 words. The parent implements and verifies. Explicit second-opinion requests are honored; routine edits and comfort checks do not trigger extra calls.
- Orchestrator: keep small shared-corpus work in one parent context. Use Luna 6/low for mechanical extraction and Luna 6/high for routine bounded reading; use Sol 6.1/high for code analysis, selectively xhigh. Astra/high is a starting point for a difficult independent branch or justified evidence review. A declared reasoning-gap escalation can consume the brief's single retry; access failures do not trigger fallback. At most six worker interactions, including retries. Reports carry evidence pointers; the parent verifies and decides. Count verification and repair when comparing costs.
- Host boundaries: read-only shell sandboxing does not automatically restrict every connector. Advisors and workers must not mutate external systems, edit project files, or recursively delegate. Respect host permissions and user scope.
These are operating defaults. A 2026-09-14 pilot using older workers (three questions over four source
files, one attempt per configuration) favored one bundled Astra session; among
delegated configurations, Luna/high plus Astra verification cost less than the
prior Luna/low, Terra/medium, Sol/high mix. This does not prove savings on larger
workloads or establish a ranking for Luna 6 and Sol 6.1. API prices, subscription usage, caching, reasoning, verification, and
retries affect actual cost; no Fable savings or trigger benchmark is claimed for
Astra. The scenarios in evals/astra-advisor.json and
evals/astra-orchestrator.json cover consult restraint, exact model selection,
host routing, worker reports, failures, and user overrides. They are evaluation
specifications, not measured GPT behavioral results.
Use Fable judgment at decisions that are costly to reverse.
The parent can run in Codex or Claude Code. It gathers the evidence, implements,
and verifies; Fable returns a compact verdict. Both consultation routes use
xhigh effort by default, with explicit user overrides respected.
| Parent host | Invocation | Model / effort |
|---|---|---|
| Codex | claude -p with a compact briefing on stdin |
claude-fable-5-1, --effort xhigh |
| Claude Code | Native fable-advisor agent |
model: fable, effort: xhigh |
A short independent review can prevent substantial rework. Keep evidence focused, cap interactions, and leave implementation with the parent. The original pattern was adapted from Anthropic's advisor tool. Actual cost depends on the selected model, reasoning, runtime context, and account billing; the older Opus/Fable benchmarks below are historical Claude Code results, not Codex measurements or a promise of current per-consult pricing.
The repository exposes .agents/skills/fable-advisor as a relative symlink to the
packaged skill. Open Codex in fable-advisor/ and ask:
Use $fable-advisor for a second opinion on this architecture before implementation.
For use in other projects, install the complete skill folder from
czlonkowski/fables, path plugins/fable-advisor/skills/fable-advisor, using Codex's
Skill Installer, or copy it into that project's .agents/skills/fable-advisor/.
It includes its CLI reference and does not require the Claude plugin to be installed.
The CLI route
uses claude -p --model claude-fable-5-1 --effort xhigh, tools disabled, safe mode,
JSON output, and no persisted session. It bounds turns, spend, and runtime, verifies
success and model metadata, and uses a fresh invocation for a justified follow-up.
Codex supplies source excerpts because the default advisor cannot open files.
Checked against Claude's programmatic-use documentation,
CLI reference, and
model configuration on 2026-09-07.
--safe-mode preserves the normal login; --bare skips subscription OAuth/keychain
access. The native agent's effort field follows the
sub-agent documentation.
| Component | What it does |
|---|---|
Skill fable-advisor |
The decision protocol: when to consult (and when not to), hard budget caps, the briefing-packet format, how to weigh the advice |
Agent fable-advisor |
A Claude Code read-only subagent with model: fable and effort: xhigh, plus a system prompt that enforces terse, committed verdicts (Verdict → Why → Risks → Would change my mind, ≤300 words) |
Recommended: add this line to AGENTS.md for Codex or CLAUDE.md for Claude
Code to make the consultation rule explicit (the trigger benchmarks below cover
Claude Code only):
Before committing to any costly-to-revert decision (architecture, DB schema, API/webhook
contracts, technology selection, production migration plans), before starting any
unattended loop/schedule/routine, or when stuck after 2+ failed fix attempts, consult
the fable-advisor skill first.
Requirements: authenticated Claude Code with the selected Fable model available. The tested Codex CLI route uses Claude Code 2.1.263; Fable 5.1 requires 2.1.255+. The native agent stays inside Claude Code; Codex invokes the CLI process instead.
We benchmarked the skill's triggering on 20 realistic queries (10 should-fire, 10 tricky near-miss negatives) across four description variants, ~200 runs on Opus. Result: zero false-fires in every variant — the skill never triggered on trivial changes, decided architectures, or questions about Fable — but recall on bare decision prompts plateaued around 50–60% regardless of description wording. The cause is structural: Claude Code consults skills only for tasks it can't handle alone, and Opus believes (correctly, in a narrow sense) that it can answer a design question itself. That belief is the exact failure mode this skill exists to counter.
If you want deterministic triggering, use the one-line CLAUDE.md setup above. Naming
it also works: prompts that mention Fable, a "second opinion", or "check with a stronger
model" trigger far more reliably (see Prompts to try below).
The gate — two questions before every consult:
- Is this decision costly to revert, or am I genuinely stuck?
- Is there a real fork in the road, with evidence to weigh?
If either is "no", the orchestrator decides on its own.
The five triggers:
- Costly-to-revert decision, before building — architecture, DB schema, API/webhook contracts, n8n workflow topology, technology selection
- Stuck escalation — 2+ genuinely different failed attempts, evidence in hand
- Plan review — a draft implementation plan embedding a costly-to-revert choice
- Pre-completion review — before declaring done on production deploys, migrations, client-facing deliverables
- Unattended-automation design — before starting a loop, schedule, or routine
(
/loop,/schedule,/goal, cron-style agents) that runs without a human watching. A loop multiplies its design flaws — a bad stop condition or interval fails on every iteration — so Fable reviews stop conditions, interval-to-change-rate match, per-iteration verification, blast radius, and cost per iteration, once, before it runs. (Read-only, easily-cancelled polling doesn't need this.)
The budget (hard rules):
- Default one consult per task, hard cap three Fable interactions
- Every consult announced to the user in one line before it happens
- One spawn + at most one reconcile follow-up per question
- Generation work (code, docs, configs, workflows) never goes to Fable
The briefing packet: decision in the first line, options with a stated leaning, hard constraints, curated evidence, ≤5 file pointers the advisor may read narrowly — and an explicit answer-format request, because advisor output is the biggest cost driver (Anthropic measured ~7× output reduction from capping, with no quality loss).
Consulting Fable on sync architecture (trigger: costly-to-revert, consult 1/3).
→ Fable verdict: nightly batch delta sync; per-document webhooks add SharePoint
subscription-renewal failure modes your one-person team can't absorb. Revisit if
freshness requirements drop below 4 hours.
In the style of the Claude Code prompt library:
plan how to migrate our API from REST to gRPC — this is hard to undo once services depend on it, so get a Fable verdict on the approach before finalizing
I'm torn between Postgres LISTEN/NOTIFY and a proper queue for job dispatch. decide, and check the decision with Fable before we build around it
review my deploy plan for Saturday's production migration and fix anything risky before I run it
I'm about to turn this on for the weekend: /schedule every 15 min, triage new support tickets and reply automatically. get a Fable review of the loop design first — stop conditions, interval, what could go wrong unattended
The skill can also trigger without Fable being named when a task hits a costly-to-revert
decision, a stuck debugging loop, or a pre-production review — but unnamed triggering is
~50–60% reliable in our benchmark. For deterministic behavior, use the one-line
CLAUDE.md setup from What's inside.
The repo ships the eval scenarios used to develop the skill (evals/): an
architecture-decision task (should consult once), a trivial-change task (should not
consult at all), and a production-migration plan review (should consult before
finalizing). Runs compare with-skill vs. no-skill orchestrators on consult discipline,
briefing compactness, and outcome quality.
Measured results (v0.1, 3 scenarios × with/without skill, Opus orchestrators):
- Consult discipline was perfect: exactly 1 consult on each costly-to-revert task, 0 on the trivial task. Briefings came in at ~511 words — well under budget — and every consult was announced to the user with the verdict relayed afterward.
- On the trivial task the skill added zero overhead (same duration as baseline, ~half the output tokens).
- The consults earned their keep: in the architecture task Fable corrected the polling cadence and contributed the failure-mode list; in the migration review it confirmed the orchestrator's 8 fixes and added 3 gaps the orchestrator had missed (window pre-staging, owner-account ordering before credential import, webhook re-registration verification).
- Trigger benchmark: ~200 runs across 4 description variants — zero false-fires, which is why the budget rules can afford to be generous about consulting.
Plan big, execute small.
Teaches a session running on Fable 5.1 (or another premium model such as Opus 5.5) to work like the coordinator in Anthropic's plan-big-execute-small cookbook: Fable plans, decomposes, and synthesizes — but never pulls bulk material into its own context. Cheap parallel workers (Sonnet/Haiku) do the token-heavy reading in their own context windows and report back distilled findings.
Most substantial tasks are two jobs in one: a little judgment and a lot of mechanical reading. On a Fable 5.1 session both bill at $10/$50 per MTok — 5× Sonnet 5.5, 10× Haiku 4.5 — unless the reading is moved. On Opus 5.5 ($4/$20) the gap to Sonnet shrinks to 2×, but it's still there. (Prices as of 2026-09.)
And there's a trap that makes naive delegation useless: subagents inherit the session
model. On a Fable session, an un-pinned Explore or general-purpose spawn runs on
Fable too — same reading, same premium rate, plus spawn overhead. The rate split only
exists when workers are explicitly pinned to a cheap model, which is the one mechanical
habit this skill enforces.
The cookbook measured the split honestly against a rigor-matched solo frontier agent: roughly 2.5× cheaper and 3× faster, with 84–98% of input tokens billed at worker rates. (Their numbers, on web research; we haven't re-measured in Claude Code yet.)
| Component | What it does |
|---|---|
Skill fable-orchestrator |
The delegation protocol: the delegate-or-read gate, five workload shapes, the worker brief format, model-tier choice, brief granularity, premise verification, and when NOT to split |
Agent worker |
A read-only reader (Read, Grep, Glob, WebFetch, WebSearch) pinned to model: sonnet (override to haiku per-spawn) whose system prompt enforces the distilled-report contract: findings with evidence pointers, never raw dumps |
Recommended: for deterministic firing, add this line to your project or global
CLAUDE.md (same reasoning as the advisor's — Claude under-consults skills for work it
believes it can do itself, and "just read everything myself" is exactly such work):
When this session runs on Fable 5.1 or another premium model and a task requires bulk
reading — sweeping code, triaging logs, reviewing documents, researching the web —
consult the fable-orchestrator skill before reading the material yourself.
Requirements: Claude Code with Sonnet/Haiku available as subagent models (the sonnet
alias resolves to Sonnet 5.5 from Claude Code 2.1.284). Designed for sessions where the
base model is Fable 5.1; the protocol applies from any premium orchestrator, including
Opus 5.5.
The gate — two questions before any bulk read:
- Is the reading mandatory and voluminous (more than a handful of files or pages)?
- Can a cheap model extract what's needed, or does the judgment live in the raw material itself?
Mandatory + voluminous + extractable → fan out. Anything else → Fable reads it itself.
The five workload shapes: codebase sweep · log triage · document review · web research · coverage verification (N facts × M sources — the shape the cookbook measured).
The worker brief: SUB-QUESTION (one line) / SCOPE (exact paths, globs, URLs) /
REPORT (shape + length cap — worker output is what enters premium context) / DON'T
(out of scope, rabbit holes). Independent briefs go out in one message, in parallel.
Fewer, bigger briefs: each spawn has a floor cost, and the cookbook found
over-splitting raised the bill.
The discipline:
haikufor mechanical sweeps,sonnetfor reading judgment, never Fable for workers- Decisions, plans, and synthesis never delegated down — workers report facts
- One premise-verification worker when the fan-out rests on an assumed list (the cookbook's own run verified 20 facts perfectly against a park list that was wrong)
- Failed worker → re-assign the brief once; don't quietly read it yourself at 5× the rate
- Reports are trusted: no re-reading what a worker read, spot-checks only narrow and only for load-bearing surprises
- The final message tells the user the shape of the run: how many workers, which models, what stayed at Fable rates
Fan-out: 6 workers (4 haiku, 2 sonnet) read ~90 files; only their reports entered
Fable context. Premise check included (service list verified against docker-compose).
audit all ~80 workflows on the n8n instance for hardcoded credentials and http:// endpoints — keep the token bill sane
go through last week of logs in /var/log/hermes and figure out why memory climbs every night around 02:00
research current pricing and rate limits for the top 5 managed vector DB providers, verified against official docs (not blog posts), and recommend one for our scale
sweep the monorepo and inventory every call to the OpenAI API with file:line, model used, and whether it goes through our retry wrapper
We ran the same 20-query trigger benchmark that shaped fable-advisor (10 realistic
bulk-reading tasks that should fire, 10 tricky near-misses that shouldn't), 3 probes
each across 5 description variants. Fable wasn't available as the CLI model at test
time, so this ran on Opus 4.8 as a proxy — read the numbers as indicative, not as
true Fable behavior.
- Zero false-fires (precision 100%) on every variant. The skill stayed quiet on all
ten near-misses — single-file reads, "explain the pattern" questions, a
second-opinion prompt that belongs to
fable-advisor, and pure generation tasks. - Recall was near-zero regardless of wording. Opus rarely consulted the skill on
genuine bulk-reading tasks, and none of four rewrites moved the needle — the original
description was kept as best. This is the same structural effect measured for
fable-advisor, and it's stronger here: Claude Code consults a skill only for work it can't easily do alone, and "read a pile of files myself" is exactly the work a capable model is confident it can do. That confidence is the failure mode this skill counters — and a stronger model (the real Fable target) tends to under-consult more, not less.
So don't rely on unnamed triggering — use the CLAUDE.md line above. It's the
deterministic mechanism; the description's job is mainly to not false-fire, which it
does well. Naming the skill, or flagging cost / "fan out" / "at Fable prices" in the
prompt, also helps.
Full eval scenarios (a coverage-shaped task graded on worker fan-out and rate split, a narrow task graded on zero overhead) are still planned. The ~2.5×/3×/84–98% figures above remain the cookbook's, on its web-research workload — not yet re-measured in Claude Code.
Built by Romuald Członkowski @aiadvisors.
MIT — © 2026 Romuald Członkowski @aiadvisors