Claude Code

Skills, and when to write one

Claude Code

A skill is a procedure loaded only when it is relevant. Write one when you keep pasting the same instructions — and not before.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

A skill is a folder with a markdown file of instructions. The agent reads only its name and description at startup, and loads the body when a task looks like a match.

That loading behaviour is the entire point, and it is what separates a skill from the alternatives.

Skill, or project instructions, or a prompt?

Project instructions (CLAUDE.md and friends) are read every session, so every line is a recurring cost. They are for facts that are always true: build commands, conventions, where things live. Keep them lean.

A skill's body is read only when relevant. That is what makes it the right home for long procedures: a 400-line checklist costs nothing on the sessions that do not need it.

Its name and description are a different matter. Those sit in a listing that loads on every turn, used or not, so a collection has a standing cost even when none of it fires. The lazy part is the body, not the skill.

A prompt is for the thing you are doing right now and will not repeat.

The rule of thumb from the documentation is direct: write a skill when you keep pasting the same instructions, checklist or multi-step procedure — or when a section of your project instructions has quietly grown from a fact into a procedure.

What makes one actually fire

The description is not documentation. It is a routing rule, and it is the only part loaded before the decision to use the skill is made.

# never fires
description: Helps with reporting

# routes reliably
description: Generates weekly metrics summaries from the analytics database.
  Use when the user asks for a weekly report, a metrics summary, or a KPI update.

Three requirements: third person (it is injected into the system prompt, and mixed point of view degrades discovery), what and when (a description with no trigger clause carries no routing information), and real trigger words — the terms someone would actually type.

If your skill never fires, the description is almost always why, and rewriting the body will not help.

There is a second cause worth knowing once you have a lot of them. The listing has a character budget, by default 1% of the model's context window, and when it overflows Claude Code shortens descriptions to fit. It drops them starting with the skills you invoke least, which can strip exactly the keywords a rarely used skill needed in order to be matched. Rarely used becomes never used, and the description you wrote is not the one the model sees.

Two defences. Put the key use case first, because each entry's description and when_to_use text is capped at 1,536 characters regardless of budget. And keep the collection small enough to fit, which is the subject of the next section.

The opposite failure has its own lever: a skill that fires when you did not want it wants a narrower description, or disable-model-invocation: true if you only ever mean to invoke it by hand.

Keep it short, for a reason

The context window is shared with everything else the agent is doing. Once a skill loads, every line competes with the actual task.

Assume the model is already competent. Do not explain what a race condition is. Keep the body under 500 lines; past that, split it into reference files — and keep references one level deep, because a file reached through another file may be read only partially, which produces confident incomplete answers.

The collection needs the same discipline, and it is measurable. /context reports what the listing costs after the budget is applied, /doctor estimates the cost and names the biggest contributors, and /skill-doctor reports how often each skill is actually used and flags the ones never invoked at all. Start with the never-invoked entries that cost the most. A skill you keep for the day you might need it is charging rent on every turn until then.

Match specificity to fragility

The taskHow prescriptiveWhat it looks like
Many valid pathsLooseProse heuristics
A preferred patternMediumA template, a parameterised script
Fragile or destructiveTightAn exact command, "do not modify"

A narrow bridge gets guardrails. An open field gets a direction. The common mistakes are opposite: creative work over-specified into a rigid checklist, and a destructive operation left to judgement.

Write the evaluations first

Before writing the body, run the task without a skill and note what actually went wrong. Turn those failures into three concrete scenarios, then write the minimum that passes them.

Skipping this is how skill collections fill with plausible skills that never trigger — written for imagined problems rather than observed ones.

Then test it in a different session from the one that wrote it. The session that wrote it already knows what you meant. What you are watching for is not "the instruction was wrong" but "the instruction was never reached".

Try this

Look at your project instructions file for a section that has become a procedure — a sequence of steps rather than a fact. Move it into a skill with a description naming the task that triggers it. Every session that does not need it stops paying for it.

What goes wrong

A description that describes. The single commonest reason a skill never fires.

Everything in project instructions. Read every session, whether relevant or not.

Explaining what the model already knows. Pure cost, every time it loads.

Writing the skill before observing the failure. You end up documenting an imagined problem.

Testing in the session that wrote it.

The collection nobody prunes. Twelve skills, three of which ever fire. The other nine are paying context rent every turn and crowding the budget that decides whether the three keep their descriptions.

How to check it worked

Start a fresh session, give it a task the skill should handle, and see whether it triggers without being named. If you have to invoke it manually every time, the description is not doing its job.

Worked examples: Anthropic's own skills repository carries the spec, a template and examples by category. For an opinionated second set, aiusage-skills is twelve skills built to this standard, with a validator that enforces the constraints above.

Sources

  1. Extend Claude with skills — Claude Code documentation Tier 1 2026-09-11
  2. Skill authoring best practices — Claude Platform documentation Tier 1 2026-08-31