GnothiGnothi
SeriesFieldsCommunityPublishing
Sign inGet started free

6. Writing a Method Down Once as a Skill

Summary

Skills, CLAUDE.md and hooks

The routine for reproducing a bug is a method, and it only matters for certain tasks. CLAUDE.md is read in full every session, so it suits standing rules. A hook is a command Claude Code runs at a set moment whether or not the model remembers. A skill is a method written down once. At startup only its name and short description are read. The model loads the full body with the Skill tool when a request matches. The rule of thumb: a procedure belongs in a skill, and anything a machine must enforce belongs in a hook.

Building reproduce-bug

A project skill is a folder under .claude/skills/ containing a SKILL.md with frontmatter. The frontmatter needs a name and a description, and the description is the only part the model sees when it decides whether to load the skill (Claude Code skills docs). This skill's description does three things: it names the result, lists the words real bug reports use, and excludes feature requests, refactoring and styling. disable-model-invocation stays off because the skill has no side effects. The body sets out five steps. The agent restates the report as exact steps, records the data, reproduces the bug twice from a clean state, and traces the code read-only. Then it stops with a four-part note: steps, data, observed result and expected result. The body stays well under the guidance of roughly five hundred lines (skill authoring best practices).

Testing and fixing the description

To test, restart the session and check /skills. Then run a positive test with an emoji bug report, which should load the skill, and a negative test with a sort-button feature request, which should not. A vague description like "Helps with problems in the app" loads the skill on feature work. A narrow one like "Use for reproduction notes" never loads it. Both are fixed by rewording the description, not the body.

The disappearing item

On the made-up report "Sometimes when I add things fast, one of them disappears," the trace points at a second submit while the first save is still in progress. In the illustrative manual check, the item vanished both times and came back on reload: it was saved, and the screen dropped it. The fix is left for a separate task. The skill gets committed, and CLAUDE.md gains one line pointing to it. Bringing in real issues and logs is left to MCP.

Early October Claude Code updates

According to the Claude Code changelog, versions 2.1.289–2.1.296 shipped these changes:

  • An onFailure: "block" option for command and HTTP hooks.
  • Hook tracking in debug mode.
  • Sonnet 5.5 cache reads halved to ten cents per million tokens, now reflected in /cost and the status bar.
  • A fix for lost final messages when a session ended.
  • Patches for several shell permission bypasses.

A skill for reproducing bugs

The last chapter ended with the type check running by itself after every edit. That hook covers a check that must happen every time. Some of the routine is a different kind of thing. It is a method, a sequence of thinking steps, and it only matters for certain tasks. Reproducing a bug is the clearest case in this course. Three chapters ago we did it by hand. We turned a vague report about pasted names into written steps, test data, an observed result and an expected result. We reproduced it twice from a clean start. Then we traced the code with editing blocked before changing anything. That method worked. But right now it lives in our memory and in an old chapter. Nothing in the repository tells a fresh session to work that way.

This chapter fixes that with a Claude Code skill. First, here is how a skill differs from the two tools we already have, because the difference decides where each piece of your routine belongs.

CLAUDE.md is the project instruction file. Claude Code reads all of it at the start of every session, so it is always there and always costs space in the context window. That makes it the right home for standing rules, such as the test command, the architectural limits and the definition of done. A hook is different. It is a command that Claude Code itself runs at a set moment, like after every edit, whether or not the model remembers. The model takes no part in deciding whether it runs. A skill is a third thing. It is a method written down once in a file, and the full text is pulled into the conversation only when a task calls for it.

That last part is the one to understand. When a session starts, Claude Code reads only the name and the short description of each skill. Each costs a small amount of context, roughly one or two hundred tokens. When you send a request, the model compares what you asked against those descriptions. If one matches, it uses an internal tool called Skill to load the full body of that skill. The rest of the time the body is not in the conversation. So a skill costs almost nothing until you need it. A long bug-reproduction checklist sitting in CLAUDE.md would cost its full size every session, even while you are styling a button.

One rule of thumb follows. If a line in CLAUDE.md stops being a fact and starts being a procedure, it probably belongs in a skill. If an instruction must be enforced by a machine rather than followed by a model, it belongs in a hook.

The gap, shown with a fresh report

Here is an illustrative scenario. It is a made-up report, not a real user. A message arrives about the shared-list app: "Sometimes when I add things fast, one of them disappears." That is all.

The repository is where we left it. It has the earlier fixes, CLAUDE.md with its definition of done, the deny list and the type-check hook. There are no skills yet. If you paste that report into a session and say "fix this," the agent will very likely start guessing. It might add a debounce to the button, or rewrite how the list reloads. Each guess changes code before anyone has watched the bug happen. Our earlier method exists to prevent that. The method is only reliable if it runs every time, and right now it runs only when someone remembers it.

Where a project skill lives

A project skill is a folder inside the repository. The path is a folder called skills inside the dot-claude folder at the repository root. Inside that is one folder per skill, named in lowercase words joined by hyphens. Inside that folder is a file called SKILL.md, in capital letters. Because it sits in the repository, it is committed with Git, and everyone who clones the project gets it. That is the same reason the hook script went into the committed settings.

We will call ours reproduce-bug. So the file is the SKILL.md inside a reproduce-bug folder, inside the skills folder.

The file has two parts. At the top is a short header block, fenced above and below by three dashes. This is called frontmatter. Below it is ordinary written instructions. Two fields in the frontmatter are required. The first is name, which is also the word you type after a slash to call the skill directly. It is lowercase with hyphens, up to sixty-four characters. The second is description, up to about a thousand characters. The description has a big job. It is the only part the model sees when deciding whether to load the skill. So it must say what the skill does and exactly when to use it.

There are optional fields too. Only one matters today. A field called disable-model-invocation, when set to true, stops Claude from ever loading the skill on its own. Then you have to call it yourself by name. That is meant for skills with side effects, like deploying or committing. Our skill only reads, runs and writes a note, so we leave automatic loading on. That is the point of this lesson.

Writing the skill

Start with the description, because it decides everything else. Here is the one we will use, spoken as written: "Turns a user's bug report into a reproduction note before any code changes. Use when the user reports something broken, wrong, missing or behaving unexpectedly in the app, or pastes a complaint from a user. Do not use for feature requests, new functionality, refactoring or styling changes."

Notice three things. It names the result, a reproduction note. It lists the kinds of words a bug report actually uses, such as broken, wrong, disappears and unexpectedly. And it names the requests that look similar but are not bugs. That last sentence is what keeps the skill out of feature work.

Then the body. It is our old method, written as numbered steps. Write it in plain sentences. In order, it tells the agent to do the following.

  1. Restate the report as exact steps a person could follow in the browser, starting from a freshly loaded app.
  2. Write down the exact data used, such as the list name and the item text, including spaces or hidden characters.
  3. Reset to a clean state, run the steps, and repeat them a second time from a clean state. Record whether the bug appeared both times.
  4. Trace the code path read-only and name the likely location. Make no edits.
  5. Stop and produce the reproduction note, then wait for the user before any fix.

The body also states the required output. The note has four labeled parts: steps, data, observed result and expected result. Below those it adds whether the bug reproduced twice and where the trace points. That output requirement matters. It gives you something concrete to check, the way the hook's exit code gave you something concrete to check.

A skill folder can hold supporting files beside SKILL.md, such as a template or a helper script, and the body can point to them. We do not need one yet. Keep the first version to one file. It also stays far below the guidance of keeping a skill body under about five hundred lines.

Testing it small first

Do what we did with the logger hook. Prove the smallest thing before trusting the bigger one.

Skills are indexed when a session starts, so if a session was already open when you created the file, restart it. Then type slash skills. Claude Code lists the available skills and where each comes from. You should see reproduce-bug listed as a project skill. If it is missing, check the folder path and the dashes around the frontmatter before anything else. You can also just ask, "What skills are available?" and the model will list what it has indexed. The slash command is the more direct check.

Next, the positive test. Send a request that is plainly a bug report, but small and not the real one: "When I add an item with only an emoji, it shows as a blank square." Watch the transcript. You should see the agent use the Skill tool to load reproduce-bug before doing anything else. Then it should go through the steps and end with a four-part note. That proves it loads and produces the right shape.

Then the negative test, which matters just as much. Send a feature request: "Add a button that sorts the list alphabetically." The skill should stay out. The agent should start on the feature in the normal way, with no Skill tool call in the transcript. A skill that fires on everything is not cheap any more. It just moves the old CLAUDE.md cost around and adds confusion.

When the description is wrong

Here is a failure worth seeing, because it is the most likely one. Suppose your first draft had a lazy description: "Helps with problems in the app."

Run the two tests again. The feature request about sorting now loads the skill too, because to the model a missing sort button is a kind of "problem in the app." The agent tries to reproduce a bug that does not exist and writes a note with no observed failure. You can tell what happened because the Skill tool call shows up on a request that had nothing broken in it.

The opposite failure is just as easy to make. Suppose the description had been "Use for reproduction notes." Now the emoji report never loads the skill. The user did not say "reproduction." They said something looked wrong. The model has nothing in the description to match against.

In both cases the fix is the same. Reword the description and leave the body alone. For loading too often, remove the vague words and add an explicit list of what the skill is not for. For never loading, add the actual phrases people use when something is broken. The version we wrote earlier does both. After each rewording, restart the session and run both tests again, the bug report and the feature request. If a skill still loads for the wrong things after a careful rewrite, there are two stronger controls. You can limit it to certain file paths, or set disable-model-invocation to true and call it yourself with a slash and its name. For this skill, a good description is enough.

Using it on the real report

Now the fresh report, for real. Paste in the message: "Sometimes when I add things fast, one of them disappears."

The skill loads. The agent restates the report as steps: open a list, type an item, press add, and right away type a second item and press add again before the first one appears. It writes the data it used, the list name and two short item names. It resets and runs the steps twice. Its trace, with no edits, points at how the add form handles a second submit while the first save is still in progress.

The note is a claim, so check it against the app. Open the browser and follow the steps exactly as written, with the same data. In our illustrative run, the second item does vanish on both tries, and a reload brings it back. That tells you something the report did not. The item was saved, and the screen dropped it. That narrows the problem a lot. Do not fix it yet. The note is the finished product of this skill, and the fix is a separate task with its own diff, test and commit.

Committing it

Commit the skill folder with a message that names the behavior, something like "add skill for reproducing bug reports before fixes." Then add one line to CLAUDE.md saying that bug reports go through the reproduce-bug skill, which produces a four-part reproduction note before any edit. Keep it to that one line. The method lives in the skill. CLAUDE.md just tells a new session the skill exists and why.

A skill is the wrong tool in two cases. If something must always hold, like never running destructive database commands, it belongs in CLAUDE.md or in the deny list, not in a file that loads only sometimes. If a check must always run, like the type check, it belongs in a hook, because a skill still depends on the model choosing to use it.

One limit remains. This skill reproduced a bug from a message we pasted in by hand. A real workflow needs the report itself, the issue, the logs and the data, brought in by the agent. That is the job of MCP, the Model Context Protocol, and other integrations, and it is the next tool in the course.

Claude Code changes in early October

Between the third and the ninth of October, Claude Code shipped versions two point one point two eight nine through two point one point two nine six. Three of the changes touch work this course is already doing.

The first affects hooks. Version two point one point two nine five added an option for command and HTTP hooks called onFailure, set to "block." With it, a hook that crashes, times out or exits with an error blocks the action instead of letting it through. Our type-check hook runs after an edit, so it cannot undo one. But any hook you later write to stop something before it happens, a PreToolUse hook, gets a firmer guarantee this way. The next action is small: when you write your first blocking hook, decide whether a broken hook should stop the work or step aside, and set this option to match.

The second also affects hooks. Version two point one point two nine six added hook tracking to debug mode. Start Claude Code with the debug flag and it logs each hook it ran, with the command, how long it took and the outcome. That is a quicker way to confirm our type-check hook is firing than adding a logger, as we did last chapter. Try it once on a single edit and look for the hook's line.

The third affects cost. In the same release, prompt cache reads on Sonnet five point five dropped from twenty cents to ten cents per million tokens, and the slash cost command and the status bar now reflect that. If you run long sessions on Sonnet five point five, your reported cost for repeated context goes down by half for that part. Check slash cost after your next long session rather than relying on older numbers.

Version two point one point two nine one also fixed a bug from the earlier release where the last messages of a conversation could be lost when a session ended. If you saw a session's final turns missing recently, updating fixes that. Several shell permission bypasses were patched across these releases too, which is another reason to stay current rather than pin an old version.