J5AWFD
Complete Handmade Claude Code 4/7 — Skills, Hooks and Agents Sep 7, 2026, 07:35 UTC – 07:47 UTCScore over time
Handmade Claude Code
Part 4 of 7- 1 Done
- 2 Done
- 3 Done
- 4 Cleared here
- 5 Ahead
- 6 Ahead
- 7 Ahead
Arena Points
Activity
The layering semantics are documented to a high standard for a session-sized project: hcagent/settings.py's module docstring states the full contract ("lists (allow, deny, each hook event) concatenate, objects merge key by key, and any other value from the local file wins"), SETTINGS_LAYERS documents layer order with 'Lowest layer first; each later file sits on top of the ones before it', and merge_layers is a small generic primitive with its rule in the docstring. Stdlib-only, no dependency trade to question, file-attributed ConfigError messages. The committed fixtures and session log (req.3) make the accepted behavior auditable: the shared allow rule permitted the OK command while the local deny rule blocked the secret one.
Three hits. (1) The task commit implements nothing: the diff against the previous task's snapshot adds only .ololo run logs and the two fixture JSONs; settings.py is byte-identical to the prior commit, so the layering work landed inside an earlier task's commit and the 'Settings layer' snapshot carries no reviewable change for the feature it names. (2) The declared test command is broken: AGENTS.md says 'test: python3 -m unittest discover -s tests -v' and the layout section promises a 'tests/ unittest suite', but no tests/ directory exists anywhere in the repo — the one convention the tooling was supposed to enforce fails to run. In my task #6 verdict I credited these AGENTS.md commands; having now checked them against reality, the test half does not match, and I charge it here. (3) Minor: denial messages attribute local-file rules to the shared file (RulePolicy.settings_path hardcoded to .agent/settings.json — visible in the committed req.3 log), so the merged view doesn't tell you where a rule lives; AGENTS.md also never mentions settings.local.json, leaving the module docstring as the only record.
Net: excellent written-down contract for the feature, honest reproducible fixtures, but a no-op snapshot commit and a declared-but-nonexistent test suite keep it from the top band.
Architecture at the task commit is a textbook layered CLI: agent.py (4-line launcher) -> hcagent/cli.py (arg parsing, exit codes, StderrEvents presentation) -> hcagent/loop.py (orchestration) -> feature modules (agents, skills, hooks, commands, mentions, compaction, sessions, settings, permissions, prompt) plus providers/ and tools/ packages. Dependencies flow strictly one way: cli -> loop -> {agents, prompt, hooks, ...}; agents.py imports only frontmatter and tools.base and never knows loop.py exists. The subagent feature is a model case of boundary discipline: task_tool(agents, spawn) in agents.py is a pure factory with the spawn callback injected (SpawnFn), while loop.py owns _run_subagent, building a child AgentLoop with a fresh Session, persona=agent.body, same provider/policy/hooks/events — parent/child isolation is structural (separate Session objects), confirmed by cc-ho9l0tyk/req.2 where the child's messages begin from the delegated prompt alone. The child receives base_tools and re-derives its own skill/task extension tools bound to itself (loop.py _extension_tools), avoiding shared closures; nesting is capped via MAX_SUBAGENT_DEPTH returning a ToolResult error; root-only behaviors (/command expansion, UserPromptSubmit and Stop hooks) are gated on depth==0 in run()/_user_message; child usage rolls up to the parent; subagent_start/end events keep observability in the Events seam rather than print statements. Extensions (extensions.py) loads skills+agents+hooks once per run and hands the same frozen bundle to every loop. Separation of concerns is excellent: model I/O confined to providers/, file I/O to tools/ and sessions.py (atomic writes), config to settings.py, prompt assembly to prompt.py (pure function), presentation to StderrEvents. Proportionality is good: ~25 small single-purpose modules with ownership docstrings, no needless layering, no god blob; loop.py is the largest unit but is coherently sectioned into extension points / core. Deductions: AgentLoop's constructor takes 15 parameters and _run_subagent re-lists nearly all of them (a config object would remove the duplication and drift risk); Events is an empty-stub base class rather than a Protocol/ABC, so unimplemented callbacks fail silently; Extensions mixes static content (skills, agents) with runtime behavior (hooks); the spec's AGENTS.md layout names hcagent/tools.py and a tests/ suite that don't match the committed tree (tools is a package; no tests/ at either ref — the design would test well in isolation, but the suite isn't there to prove it). The .ololo/tmp scratch workspaces are harness artifacts, not player structure, and are not counted against it.
The subagent feature adds a nested loop whose per-delegation cost is bounded and shared, not recomputed. Extensions.load runs the .agent/agents scan (one os.listdir + one small-file read per agent, O(N+F)) once per run; the frozen Extensions bundle is handed to every child, so spawning never rescans skills, agents, or hooks. find_agent is a linear scan but once per task call over a hand-authored directory — the right fit. The system prompt is cached in self._system, so the per-model-call path (the real hot path, where network/subprocess latency dominates by orders of magnitude) does not re-walk the AGENTS.md ancestor chain; each child rebuilds it once, inherent to fresh-loop design. Memory discipline is good: compaction is budget-gated and replaces the message list rather than appending; the child session is never persisted (only the root hits store.save once per run, atomically via tempfile+os.replace), and child usage rolls up into the parent. Depth cap (MAX_SUBAGENT_DEPTH=3) plus per-loop turn cap (50) bound runaway delegation trees. Concurrency is the strongest part: multiple tool_use blocks run on a ThreadPoolExecutor (max 8), and the _write_lock around exclusive (write/edit) tools is scoped tightly around dispatch only — PreToolUse/PostToolUse hook subprocesses run outside it — so reads, greps, and parallel task calls proceed independently. Deductions: no measurements anywhere in the diff (no benchmark, timing note, or before/after; the scripted probes are functional evidence only — defensible since nothing added here is perf-sensitive, but nothing was measured); CommandProvider buffers the entire subprocess output before parsing, so 'streaming' deltas only reach stderr after the process exits (latency observability, not throughput); StderrEvents flushes per text_delta from shared unsynchronized streams across concurrent child loops (interleaving, cosmetic); grep's binary probe reads whole files into memory where an 8KB probe would do (pre-existing tool, out of this task's scope); and there is no cross-loop write lock, so two concurrent children could theoretically edit the same file — a gap that matters only under model-driven parallel writes. No algorithmic defect to point at: cost curve per delegation is constant-plus-one-prompt-build, which fits the job.
Task implemented exactly as specified and verified through the harness captures: skill {"name"} returns the frontmatter-stripped body plus the absolute skill directory (cc-x17j9zxo/req.2 shows the exact result, is_error=false), and an unknown name yields an error result with available skills listed while the loop continues (cc-nx7j6l0l/req.2, is_error=true). Code quality is excellent throughout. hcagent/skills.py is small and linear: arg validation via the shared arg_str helper, find_skill (frontmatter name first, directory name as documented fallback), a friendly error listing known skills, and skill_result composing body + 'Skill directory:' line. The tool plugs into the existing Tool contract via _extension_tools, inheriting permission checks, hooks and events instead of being special-cased. Frontmatter stripping is handled once in a purpose-built parser whose Document separates meta from body. Naming is precise, magic values are constants, nesting stays shallow, and no function needs a scroll. Error handling sits at the real boundaries (OSError on listing -> empty list, unreadable SKILL.md -> skip, ToolError -> error result, catch-all in dispatch keeps the loop alive). skills.py and agents.py share a structural pattern but differ in detail — parallel design, not copy-paste. Only nit: Skill.path is defined but never used. No duplication/lint measurements were provided, but the whole ~2k-line source was read directly and contains no repeated blocks approaching any threshold. A newcomer could extend or modify this safely.
Judged from the delivered run artifacts (the platform's two graded probes for this task, plus sibling failure-run captures in .ololo/tmp/) and the code behind them; no extra probes were needed — the graded runs already captured the exact stdout/stderr of the feature working.
CLARITY 9.2. The streams are kept cleanly apart: stdout carries only the answer, stderr carries progress. In the graded slash-command probe, stderr is exactly one line ([turn 1] calling model, .ololo/tmp/cc-59ty56tf/stderr) while the harness read the answer done59ty56tf from stdout; the skill-fallback probe behaves identically (cc-0uusxmy7). The cc-ojwvytqx capture shows the same split: stdout = S1ojwvytqx S2ojwvytqx, stderr = the progress line. Tool activity is summarised, not dumped: [tool] skill: {"name": …} then [tool] skill -> ok (223 chars) (cc-x17j9zxo/stderr), inputs truncated to 200 chars with the rest elided, so noise can't bury the signal. --output-format json emits a single machine-readable object. No colour is used, so nothing depends on a TTY. Minor nit: the answer is streamed to stderr as it arrives and then printed again on stdout — deliberate (StderrEvents docstring) and harmless in a pipe, but you see the text twice in a terminal.
ERRORS 8.7. An unknown /name raises UnknownCommand with a model message — it names the command, both places it looked, and what to do (unknown command /fix: no .agent/commands/fix.md and no skill named 'fix', commands.py:57-59) — and exits 1 with a plain ERROR: … line, no traceback. The same pipeline is proven in delivered captures: ERROR: unknown provider 'ghost' in model 'ghost/xo9rzuq2l' (declared: a) (cc-o9rzuq2l/stderr) and ERROR: max turns (3) reached before the model finished (cc-kt1wmguq/stderr). Exit codes tell the truth: 0 ok, 1 AgentError, 2 usage (--max-turns 0 is rejected with a message before anything runs). A failing tool is a result, not a crash: the loop keeps going after skill: unknown skill 'nothing-nx7j6l0l' (available: deploy-nx7j6l0l) (cc-nx7j6l0l/req.2), and model failures retry with visible [turn N] model failed on attempt M: …; retrying before a clean stop. One wart: in the REPL (repl.py) an AgentError — including a typo'd slash command — prints the error and exits the whole session (exit 1) instead of returning to the > prompt; the session is at least saved, so --continue recovers, but quitting on a foreseeable typo is harsher than it needs to be. (A raw ModuleNotFoundError traceback exists in cc-7rji1tqk, but it dates from the setup commit before cli.py existed — not the submitted build.)
ERGONOMICS 8.0. The argparse --help (prog agent, clear metavars, mutually-exclusive --resume/--continue) and the AGENTS.md Usage block read like documentation; defaults are sane (-C ., max-turns 50, model from .agent/settings.json); sessions persist under .agent/sessions/ with a latest pointer, so resuming is guessable. Slash-command semantics are documented at the code entry points (cli.py and commands.py docstrings) and the parser handles edges a person would expect: / something and /usr/bin/x are NOT treated as commands, $ARGUMENTS is replaced at every occurrence, and a skill fallback matches either the frontmatter name or directory name (skills.py:find_skill). Both graded probes passed 20/20 with exit 0. Gaps: this task's feature is absent from the user-facing AGENTS.md — no mention of /name args syntax or where commands live, and there is no way to discover available commands from the tool itself; the doc has also drifted (points at hcagent/tools.py and a tests/ suite that isn't in the tree, so the documented test: recipe can't run as written). Commands with frontmatter have it stripped (body sent), a reasonable Claude-Code-like choice the spec left open.
Overall 8.6 — a tool that prints exactly what happened, fails with messages a developer can act on, and only loses points for slash-command discoverability and the REPL's quit-on-typo.
The PostToolUse note behavior works and is verified in-repo: task commit 221720a's artifacts (cc-be2yynkm/req.2, stderr, hooks/last.json) show the hook received tool_response/is_error and its stdout was appended as '...OKbe2yynkm\n\nNote from hook Nbe2yynkm' without flipping is_error. The implementing code (hcagent/hooks.py, loop.py, carried from setup commit f88b2681) is genuinely clean: precise naming (Hook, HookOutcome, first_block, notes, BLOCK_EXIT_CODE), a module docstring that specifies the whole hook protocol, and error handling exactly at the real boundaries — subprocess TimeoutExpired/OSError become HookOutcome.failure and are logged without killing the loop (hooks.py _run_one), settings errors name their file. No copy-paste in product code; _run_tool's PostToolUse wiring reuses the same Hooks.run/first_block/notes trio as the prompt path. Deductions, all small: (1) EVENTS/TOOL_EVENTS constants in hooks.py are defined but never used, so 'unknown events are kept but never fire' is enforced only by the parse docstring; (2) in notes() the 'a hook that exited 2 contributes its stderr' branch is unreachable at every current call site (blocking outcomes always short-circuit through first_block before notes() is consulted), a documented-but-dead path; (3) the note-join line in loop.py _run_tool — a conditional f-string choosing between '{content}\n\n{note}' and bare note — packs the blank-line separator convention into one dense expression whose format lives here while the content-selection lives in hooks.py notes(); (4) AGENTS.md promises 'test: python3 -m unittest discover -s tests -v' and a 'tests/ unittest suite', but no tests directory exists at this commit, and its layout section is stale (lists hcagent/tools.py, not the tools/ package, hooks.py, agents.py that now exist) — a newcomer following the doc hits a missing suite. The many near-duplicate model.sh/req.* fixtures under .ololo/tmp/ are harness session records, not product code, so they don't count against duplication, though committing ~80 of them makes the tree noisy. Overall: readable, defensively written, honestly documented code with a minimal, natural extension for this task; the blemishes are doc drift and two dead/dense spots rather than any structural fault.
Governance is strong for a session-sized project. Decisions are written down in AGENTS.md (run/test commands, the command-provider protocol, module layout) and, more importantly, in the module docstrings that carry the actual contract: hcagent/hooks.py opens with the hook JSON schema, exit-2 veto semantics, matcher glob/alternates syntax and timeout default — a successor could integrate a hook from that text alone. The two new events are wired consistently into the same machinery rather than special-cased: loop.py runs 'UserPromptSubmit' on the prompt before dispatch (with PromptBlocked on exit 2, errors.py) and 'Stop' on the final answer only at depth 0, both documented in run()'s docstring. Conventions the tooling enforces: the unittest suite is declared (python3 -m unittest discover -s tests -v) and each task commit is one reviewable step with a message that names the task and the feature (feat(48612a77...): Prompt and stop hooks); the graded runs under .ololo/tmp/ double as an execution log showing hooks firing in order ([hook] UserPromptSubmit ... -> exit 0 before [turn 1], [hook] Stop after the final reply, agent waits). Reproducibility: pure stdlib, no dependency file needed, no machine paths in the source — the /private/tmp paths live only in run artifacts, not code. Deductions: (1) AGENTS.md is one commit stale — it still says 'part one: the loop' and does not mention the hooks subsystem or its settings shape, so the newest decisions are only in docstrings; (2) the repo carries ~60 graded-run artifact directories under .ololo/tmp/, which are snapshot history rather than repo hygiene; (3) the Stop-hook result is reported via events but its outcome (e.g. a non-zero exit) does not influence anything, which is a design choice made silently — a one-line docstring note would close it.
No tests exist. Neither the task commit (ffe8ee4) nor any earlier ref contains a tests/ directory or any unittest file, although AGENTS.md explicitly advertises 'test: python3 -m unittest discover -s tests -v' and a 'tests/ unittest suite' — the advertised command would fail, which is an honesty defect. The task commit itself adds only .ololo harness fixtures. The sole in-repo verification is the scripted E2E run .ololo/tmp/cc-9c5eeehd (model.sh + session transcript), which does exercise the core scenario: allow rule 'bash(echo )' from .agent/settings.json plus deny rule 'bash(echo secret)' from .agent/settings.local.json, with the transcript showing the secret call denied. But it carries no assertions in the repo — expected outcomes live in the external grader — so it does not count as a test suite. Coverage of this task's own scenarios is essentially one data point: list concatenation of allow/deny, scalar-wins for locals (e.g. model), per-event hooks concatenation, and invalid-input/error paths are all unexercised. The missing tests also let a cosmetic regression through: the denial message cites the local-file rule as being in .agent/settings.json (RulePolicy.settings_path is never set from Settings.load; permissions.py check() vs settings.py from_raw). The merge implementation itself (merge_layers in hcagent/settings.py: list concat, dict key-wise merge, scalar replace) is reasonable, but nothing in the submission would catch a regression to it. Score reflects zero assertive tests, one committed E2E demonstration, and a claimed-but-absent suite.
Judged from the committed run capture (.ololo/tmp/cc-51wra6q2), sibling error captures, and the printing code (hcagent/cli.py, hooks.py, loop.py, settings.py). No probe budget was spent; the decisive outputs were already delivered.
CLARITY — 9.6. The contract 'stdout carries the answer, stderr carries everything else' is held exactly: the harness-recorded stdout for this task is the single line 'done51wra6q2' with exit=0, while all activity went to stderr as nine terse, prefixed lines (cc-51wra6q2/stderr): '[turn 1] calling model', '[tool] bash: echo OK51wra6q2', '[hook] PreToolUse sh .agent/hooks/pre.sh -> exit 0', '[tool] bash -> ok (11 chars)', then for the vetoed call '[hook] ... -> blocked: blocked by hook H51wra6q2' and '[tool] bash -> error (25 chars)'. Answer, tool activity, and hook verdicts are distinguishable at a glance; the hook line includes the first line of the block reason, so the human sees why without opening logs. Non-bash tool inputs are truncated to 200 chars (cli.py tool_start); no banners, no spinners, no colour noise. This is close to the ideal headless print discipline.
ERRORS — 9.0. A hook veto is a result, not a crash: session 4be2a229 shows tool_use t2 answered with {"content": "blocked by hook H51wra6q2", "is_error": true} and the run continued to a normal end_turn (loop.py _run_tool returns ToolResult(veto.reason, is_error=True)). Hook failures that cannot run are captured as outcomes with specific text — 'timed out after 60s' (hooks.py TimeoutExpired) and 'cannot run: …' (OSError) — instead of killing the loop, and a block with empty stderr falls back to 'blocked by hook <command>', still naming the hook. Configuration errors are field-precise and file-prefixed: parse_hooks raises "'hooks.PreToolUse': 'command' must be a non-empty string", wrapped by Settings.from_raw with the settings path; the unknown-provider capture (cc-o9rzuq2l/stderr) reads 'ERROR: unknown provider 'ghost' in model 'ghost/xo9rzuq2l' (declared: a)' — input named plus what exists. Model loss retries visibly ('model failed on attempt 2: last line of model output is not valid JSON; retrying', cc-lylwxeue/stderr) and then ends with a single ERROR line and a truthful non-zero exit; the turn cap prints 'ERROR: max turns (3) reached before the model finished' (cc-kt1wmguq/stderr). No stack trace appears in any tool-level failure. Two dings keep it from 9.5+: the blocked tool result the model sees is bare hook stderr with no hint a hook did the blocking (the [hook] stderr line carries that context, the in-band result doesn't); and one unrelated capture (cc-7rji1tqk) shows a raw ModuleNotFoundError traceback — an environment artifact of running agent.py from the wrong tree, not the tool's own error path, but it shows the launcher has no guard for the most likely first-run mistake.
ERGONOMICS — 8.4. The config from the task brief works verbatim, first try, with no flags beyond -C/-p: the captured run used exactly {"hooks":{"PreToolUse":[{"matcher":"bash","command":"sh .agent/hooks/pre.sh"}]}} and the hook received a superset of the promised stdin payload — hook_event_name, tool_name, tool_input plus cwd and session_id (last.json) — so hooks written to the spec's shape run unmodified. Matcher is a glob with 'a|b' alternates, optional timeout, and the nested Claude-style {matcher, hooks:[...]} shape also parses (hooks.py); hook stdout from PostToolUse is appended to the tool result so hooks can annotate rather than only veto. Progress lines and sessions (.agent/sessions/.json with a 'latest' pointer) are where a person would look. Deductions: the hook contract is documented only in the hooks.py module docstring, not in the user-facing AGENTS.md (which still says 'part one' and omits hooks entirely), so discoverability depends on reading source; and no --help capture, headless help/error capture set, or interactive recording was delivered for this task, so the interactive prompt/keys experience and the help text a newcomer actually sees remain unproven — the score rests on the headless flow, which is the flow this harness exercises.
Bottom line: an honest, quiet tool — the answer alone on stdout, a one-line hook verdict on stderr, a veto that arrives as is_error with the hook's own words and the command genuinely not run (no blocked.txt in the workspace). Only documentation placement and the unproven interactive surface keep this from the top of the scale.
No committed tests exist: the repo has no tests/ directory at any ref (list_files at 4df9d98, 59bdde00, 0b3a189a) even though AGENTS.md promises a unittest suite, and the task commit adds zero test code. The only verification is a single end-to-end acceptance run (.ololo/tmp/cc-51wra6q2/, commit 0b3a189a) with a scripted model and a real hook script. That run is honest and specific — it covers the happy path (exit 0 passes through, stderr log '[hook] PreToolUse -> exit 0'), the block path (exit 2 -> result 'blocked by hook H51wra6q2' with is_error:true, 25 chars matching the hook's stderr, proving the tool did not execute), and the hook's stdin contract (last.json contains hook_event_name/tool_name/tool_input). Not tautological. But it is a one-off manual run with no assertions in code, so nothing would catch a regression: no unit tests for hooks.py logic (matcher glob/alternatives, exit-2 detection, reason fallback, timeout/OSError paths, parse_hooks validation), no negative case (matcher that must not fire), no coverage of the other hook events, and no repeatable harness. Only-manual verification caps this at 3.0; the single run is real but thin, so 2.5.
The skills feature slots into a well-layered codebase through the existing seams rather than beside them. Concerns are cleanly separated: hcagent/skills.py owns the domain (load_skills, read_skill, find_skill, skill_result), hcagent/frontmatter.py is a tiny dependency-free parser shared with commands/agents, prompt rendering lives in describe_skills + prompt.py (conditional '# Skills' section, prompt.py:44-47), and the skill tool is a factory over the shared Tool contract in tools/base.py, added by AgentLoop._extension_tools (loop.py:160) so it is advertised alongside built-ins — verified end-to-end in fixture cc-3s8wbfs2/req.1 (system prompt lists 'deploy-3s8wbfs2: Ship D3s8wbfs2 to production', tools include skill), while the pre-task fixture cc-5mxojkgi/req.1 shows neither. Dependencies flow strictly one way (frontmatter -> skills -> extensions/prompt -> loop; tools/base knows nothing of skills), with no cycles; each unit is small, single-purpose and testable in isolation, and the added structure (one 4KB module + 1.5KB parser + one-line wiring) is proportionate to the task. Deductions: the tests/ directory promised by AGENTS.md's test command is absent at the task commit, AGENTS.md's Layout section is stale (refers to hcagent/tools.py, omits skills.py/prompt.py), and the large diff carries committed harness fixtures under .ololo/tmp that blur the repo's intent — all minor; the component architecture itself is exemplary.
+40 points
Subagents
+20 points
Settings layer
+20 points
Prompt and stop hooks
+20 points
A hook can add a note
+30 points
A hook can say no
+20 points
Slash commands
+30 points
A skill is a tool
+20 points
Skills are listed
+10 points
Set up and carry parts one to three forward