Which phases a question needs and in what order, and the single budget they share. It dispatches and produces no findings of its own.
- classify
- depth-tier
- chain-selection
- budget
- assembly
Agent Skills / research set
Six skills over the research chain — scoping the question, building the corpus, grading each source, synthesising a claim ledger, and writing it up with the hedges preserved exactly as the sources stated them.
A skill arrives in three stages and each costs something different. Almost every design rule in this repository follows from that asymmetry rather than from taste.
Each owns one kind of work and states it in its own Owns section.
The tags under it are the facets recorded in
research-registry/capabilities.yaml, which is also what the fixtures route
against.
Which phases a question needs and in what order, and the single budget they share. It dispatches and produces no findings of its own.
The brief everything downstream serves — what is being asked, what an answer would look like, what would falsify it, and when to stop looking. Nothing is retrieved here.
The corpus and everything known about where it came from. This is the only skill that reaches the open web, which is what makes provenance checkable at all downstream.
What each source is worth, for a specific claim. It recommends confidence and never assigns it — a label set here becomes an unremovable floor downstream.
The claim ledger and the answer it supports. On any chain that includes it, this is the only skill that assigns a confidence label.
The deliverable, and nothing that is not already in the ledger. No new claims enter at write-up time.
Boundaries live in one file and never in a description. A description that named its neighbours would spend listing budget advertising them, and adding a skill would mean editing every other one — so the cost of an addition would stop being O(1).
| Skill | Does not do | Goes to |
|---|---|---|
research-route | a request whose phase is already obvious | that phase directly |
research-route | a question answerable from what is already known | answering it |
research-scope | retrieving anything | research-source |
research-scope | deciding what a retrieved source is worth | research-appraise |
research-source | grading what was retrieved | research-appraise |
research-source | deciding what to look for | research-scope |
research-appraise | assigning the confidence label | research-synthesize |
research-appraise | finding more sources to fill a gap | research-source |
research-synthesize | writing the deliverable | research-report |
research-synthesize | re-grading a source | research-appraise |
research-report | reaching a conclusion the ledger does not carry | research-synthesize |
research-report | filling a gap the draft exposes | research-source |
A chain of names expresses linear work only. Where a stage repeats until a condition holds, the entry carries the condition, the judge, and a hard cycle limit — without all three, “until it looks right” has no stopping rule and the loop ends when somebody gets tired.
| Route | Pattern | When | Chain | Condition |
|---|---|---|---|---|
deep | linear | contested, consequential, and the answer must survive challenge | research-scope → research-source → research-appraise → research-synthesize → research-report | gate · the brief carries both round caps before any retrieval is dispatched |
standard | linear | a real question with a decision behind it | research-scope → research-source → research-appraise → research-synthesize | — |
quick | linear | a settled fact that still needs a source | research-source → research-appraise → research-report | gate · the report says no synthesis pass ran, and caps confidence accordingly |
vet | linear | somebody handed you a claim and you need to know if it holds | research-source → research-appraise → research-report | — |
frame-only | report-only | nobody can say what the question is yet | research-scope | stops at · the brief: question, answer shape, sub-questions, bar, stopping rule, prior, falsification. Nothing is retrieved |
corpus-only | report-only | gather the material, judgement comes later | research-source | stops at · the manifest and the search log. Nothing is graded |
fill-the-gap | loop | synthesis found a claim nothing supports | research-source → research-appraise | oracle · every load-bearing claim reaches its evidence bar, or is recorded as unsupported |
Eight sets share one validator and one contract shape. Each declares exactly one signature mechanism, and that declaration is the whole difference — which is what keeps 8 copies of the same rules from becoming 8 dialects.
_research/REACH.mdA source is only evidence for what it can actually reach. Every claim carries how far the sourcing got — whether the primary was opened, whether a disconfirming search was run — so a conclusion built entirely from secondary coverage says so on its face.
A rule states the mechanism inside every skill through a delivered block. A skill that owes it must also name its own half in its own words — a rule stated everywhere and owned nowhere is a ritual.
Reporting completion without meeting this is reporting a wish. The
vocabulary is fixed in research-registry/harness.yaml and defined in
_research/CONTRACT.md; the validator checks the contract actually
defines every word it declares.
| Axis | Vocabulary |
|---|---|
| Evidence grades | P1 P2 P3 P4 P5 |
| Residual classes | BLOCKED OUT-OF-SCOPE DEFERRED UNSUPPORTED |
| Statuses | DONE PARTIAL BLOCKED |
| Sizing tiers | quick standard deep |
| Route patterns | linear loop report-only |
| Reach | primary one-hop chain blocked no-primary |
DONE.A class is declared once in the registry and the validator checks each
SKILL.md front matter matches it exactly. Where a CLI does not enforce tool
grants, the Never lines are discipline and nothing more — that limit is stated
rather than papered over.
| Class | Tools | Writes | Skills |
|---|---|---|---|
route | Read, Grep, Glob, Skill |
no | 1 |
retrieve | Read, Grep, Glob, WebSearch, WebFetch, Write |
yes | 1 |
doc-write | Read, Grep, Glob, Write, Edit |
yes | 4 |
Every threshold is declared in one place and read from there by
research-tools/validate.py. A number written twice is a number that drifts.
Raising a limit is the last resort — merging, deleting, compressing and relocating come
first.
skill_md_lines | 155 |
description_chars | 200 |
playbook_lines | 300 |
playbooks_per_skill | 8 |
shared_file_lines | 136 |
shared_lines_total | 780 |
skills_max | 8 |
routes_max | 12 |
route_stages_max | 6 |
repo_md_lines_total | 5000 |
| skills | 6 / 8 |
| routes | 7 / 12 |
| shared contract lines | 741 / 780 |
| repository markdown | 3,469 / 5,000 |
| static rules | 38 |
Adding a rule means adding a deliberate violation and watching it fail. A check only ever seen passing may be checking nothing.
Each research-* directory is symlinked individually, so a
skills directory keeps whatever else it already carries and a name already taken by a real
directory is skipped rather than overwritten.
make link # into ~/.claude/skills make link CLAUDE_DIR=.claude/skills make check # the rules, then proof the rules still fire make render # after editing a delivered block make hooks # run the rules on every commit
_research/ is read on a minority of launches, so the operative part
is copied verbatim into every skill and a rule fails if any copy has drifted.