Skip to content

Skill Vault · Build and check

Find the cause of a bug

Reproduce a failure, test the explanation, then check the fix.

A failing behavior and a project where your agent can run checks.

Skill name /diagnosing-bugs

Your next step

Try it for yourself

You’ll need: A failing behavior and a project where your agent can run checks.

Paste into your coding agent with the relevant project open.

This is a one-off starter for the approach. Installing the full Skill adds its complete instructions.

Ready to copy
Diagnose this bug without guessing. First create or identify one red feedback loop for the failure, minimize it, rank hypotheses, instrument the strongest one, then fix and add a regression test. Redact secrets from all evidence.

What happens nextEvidence identifies the cause, and a regression test checks the fix.

Use it again

Add the full Skill.

The starter lets you try the approach. Installation adds the complete instructions to your AI coding tool.

Copy the setup instructionsFor Codex or Claude Code on your computer

Your next step

Ask your agent to help you install it

You’ll need: Node.js with npx, Git, and the agent you choose. A project folder where you want the Skill available.

Paste this into Codex or Claude Code with your project open. Your agent will help you review and install the package.

Review files and permissions before accepting an install. Adding a Skill does not run it.

Ready to copy
Help me install /diagnosing-bugs from MattEspo23/skills v0.1.2 in this project.

Review the package instructions and supporting files first. Include these Skills: diagnosing-bugs.

Use the command for the agent I am using:
Add to Codex:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill diagnosing-bugs --agent codex --copy

Add to Claude Code:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill diagnosing-bugs --agent claude-code --copy

Show me where the files will go before making changes. Preserve existing Skills and customizations. Stop if the release, supporting Skills, or target agent cannot be verified. Do not run the Skill or change external services during setup.

After installation, explain how I can use /diagnosing-bugs and what access it needs.

What happens nextYour agent should review and add /diagnosing-bugs, then explain how to use /diagnosing-bugs. Stop if a dependency or version cannot be verified.

No additional supporting Skill is required by this package. Source version: v0.1.2.

Prefer a terminal command?

Add to Codex

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill diagnosing-bugs --agent codex --copy

Add to Claude Code

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill diagnosing-bugs --agent claude-code --copy

Inspect the source on GitHub

Installing copies the instructions into your chosen agent. It does not run the Skill or configure project tools.

Keep the package’s supporting files, LICENSE and NOTICE together with the Skill instructions.

Full notes & source materialThe complete original text, examples and reference details.

These are the complete original notes. Planned videos and services mentioned here may not be available yet; the actions above reflect what you can use on this site now.

Diagnose a hard bug through a tight failing loop before proposing or implementing the fix.

Watch

An EspoAI verification network connecting evidence, checks, and system health.

Skill Vault video planned · Video planned

Install

Public release v0.1.2 — install the latest source or pin the tested release.

Install latest

Command
npx skills@latest add MattEspo23/skills --skill diagnosing-bugs

Reproducible install

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill diagnosing-bugs

Clean installs are verified for Codex and Claude Code. The installer copies editable files into the selected agent; review every Skill before giving it tool access.

This is the public, editable adapted baseline. Its deeper Espo-specific revision and Skill Vault video are planned; the released source can be installed now.

What it does

diagnosing-bugs runs a six-phase diagnosis on a hard bug or a performance regression: build a repro, minimise it, rank hypotheses, instrument, fix with a regression test, clean up.

It will not let the agent form a theory until a tight feedback loop exists — one named command, already run once, that goes red on this bug and green when it is fixed. The default behaviour of a coding agent handed a bug report is to read code and guess; this skill blocks that. If no red-capable command exists, there is no Phase 2. That single gate is what the skill is for. Everything after it — bisection, hypothesis-testing, instrumentation — is mechanical once the signal exists.

When to reach for it

Type /diagnosing-bugs, or the agent reaches for it on its own when a task fits — it is model-invoked, and fires on "diagnose" / "debug this" or on a report that something is broken, throwing, failing, or slow.

Reach for it on the hard ones: a bug that resists a first look, an intermittent flake, a regression that crept in between two known-good states. It is heavy by design, and the wrong tool for a question you want answered in one message.

Your situationWhere to go
A specific defect you can describe as a symptomThis skill
A slow endpoint or a timing regression with a known before-and-afterThis skill — it has a performance branch (measure a baseline, then bisect)
"Where are the bottlenecks in this codebase?" — no specific symptomNot this skill. It diagnoses one known failure, it does not audit
A raw bug report from someone else, not yet confirmed or written uptriage first
Throwaway code to answer a design question, not chase a defectprototype
Building a planned behaviour test-firsttdd
No good seam exists to lock the bug downimprove-codebase-architecture — this skill hands off there itself

The tight loop is the skill

Phase 1 gets disproportionate effort because it is the only phase that is hard. The skill gives a ladder of ways to construct the loop, roughly in order of preference:

  1. A failing test at whatever seam reaches the bug.
  2. A curl or HTTP script against a running dev server.
  3. A CLI invocation with a fixture input, diffed against a known-good snapshot.
  4. A headless browser script asserting on DOM, console, or network.
  5. A replayed capture — a saved request, payload, or event log, run through the code path in isolation.
  6. A throwaway harness: a minimal subset of the system, one function call.
  7. A property or fuzz loop, for "sometimes wrong output".
  8. A bisection harness you can hand to git bisect run.
  9. A differential loop — same input, old version against new.
  10. A human-in-the-loop bash script, last resort. The skill ships scripts/hitl-loop.template.sh for this: the agent runs the script, you follow prompts in your terminal, and your answers come back as parseable output.

A loop is not the goal. Tight is: fast (seconds), deterministic (same verdict every run), sharp (asserts your exact symptom, not "didn't crash"), and agent-runnable unattended. A 30-second flaky loop is barely better than none. For a bug that only shows up sometimes, the target is not a clean repro but a higher reproduction rate — loop the trigger, parallelise, add stress, inject sleeps, until the flake rate is high enough to debug against.

When it genuinely cannot build one, it is instructed to stop and say so, list what it tried, and ask you for environment access, a captured artifact, or permission to add temporary instrumentation. It should not proceed to hypothesise anyway.

The gates between phases

The phases are gates, not a checklist. Each one refuses to open until something specific is true.

GateWhat has to be true
Into Phase 2A named command, already run and pasted with its output, that can go red on this bug
Into Phase 3The repro is reproduced and minimised — every remaining element is load-bearing
Into Phase 43–5 ranked, falsifiable hypotheses exist, each stating its prediction, shown to you before any is tested
Into Phase 5Probes map to a specific prediction, one variable at a time, every debug log tagged [DEBUG-a4f2]-style so cleanup is one grep
DoneOriginal repro no longer reproduces, instrumentation gone, and the hypothesis that turned out correct is written into the commit message

Phase 5 has an escape hatch worth knowing about. The regression test is written before the fix, but only if a correct seam exists for it — one where the test exercises the real bug pattern as it occurs at the call site. Where the only available seam is too shallow, the skill is told to say so rather than write a test that gives false confidence. That absence is itself the finding, and it is what routes the post-mortem to improve-codebase-architecture.

Common questions

It fires on quick questions where I just wanted a direct answer. This is the most-reported problem with the skill, and it is real. On GPT-5.6-Sol especially, users report it triggering on a plain description of a problem: "the model triggers the rather formal diagnosing-bugs skill instead. It then goes on to construct a reproduction scenario — often building a mock scenario with limited value — before giving me a response or suggestion. This results in considerable reply delays." Four separate people reported the same shape on issue #578. The accepted fix is to start with a lighter approach and graduate to the heavier one only where the problem warrants it, but that change has not landed. The skill is calibrated against Claude Code's invocation behaviour; a model with a lower activation threshold over-fires it. Until it is graduated, the practical fix is to say what you want ("just answer this, don't diagnose") or to disable model invocation for it in your harness.

Can I point it at a codebase and ask where the performance problems are? No. It diagnoses one failure you can already name. Its performance branch is for a regression with a symptom — establish a baseline measurement, then bisect, measure first and fix second — not for a proactive sweep. A skill for the proactive version was proposed and closed; there is currently no skill for it.

Does it stop and ask me before it writes the fix? No. Only Phase 3 has a human checkpoint — the ranked hypothesis list is shown to you before any is tested, and it proceeds on its own ranking if you are away. There is no gate between instrumentation and the fix, so the agent can start writing code before you have agreed with its root cause. Issue #124 asks for that gate and is still open. If you want it, say so when you invoke the skill.

I already ran /triage on this bug report. Is this the same work again? Partly, and neither skill admits it. As one reader put it: "Triage's step 3 is essentially a shallow, bounded instance of diagnosing-bugs Phase 1–2, but neither file mentions the other." Triage does a bounded "is this actually a bug, and what is the surface" pass; this skill does the thorough version. Running triage first is not wasted — its verification often gives you most of Phase 1's raw material — but expect to redo it properly here, and expect no cross-reference to tell you that.

Will the repro output it pastes leak secrets? It might. The skill asks the agent to paste the invocation and its output, and to request artifacts like HAR files, log dumps, and core dumps. None of those are sanitised by instruction. Issue #674 raises exactly this — credentials, tokens, cookies, and personal data riding along into a chat, an issue, or a PR — and proposes a redaction guardrail. It is open and unimplemented. Treat redaction as your job for now, particularly before the output goes anywhere public.

My security scanner flagged this skill as high risk. Snyk flags it, and the flag is a false positive. It is the only skill in the set that ships an executable shell script (hitl-loop.template.sh) alongside instructions to run it and to curl a dev server. Shipped .sh plus run-it instructions plus outbound HTTP is enough to trip a static scanner. The script itself is about 30 lines of read -r -p prompts that pause for human input. The scanner is rating the capability surface, not a proven exploit.

What happened to /diagnose? Renamed to /diagnosing-bugs in v1.0.0. The old name no longer exists. Anything of yours that chains /diagnose — a wrapper skill, a saved prompt — needs updating.

It's working if

  • It shows you a command and its red output before it offers a single theory. If theory arrives first, the skill is not running.
  • The failure it reproduces is the one you reported, not a nearby one it found on the way.
  • It shrinks the repro before it starts guessing, and can tell you why each remaining piece is load-bearing.
  • You are shown a ranked list of 3–5 hypotheses, each with a prediction you could falsify, before any of them is tested.
  • Every debug log it adds carries a tag like [DEBUG-a4f2], and a grep for that tag comes back empty when it declares done.
  • The commit or PR message names which hypothesis was right.
  • When it cannot lock the bug down with a test, it says so plainly instead of writing a shallow one.

Where it fits

diagnosing-bugs is a reach-for-it-anytime standalone. You drop into it when something is broken and drop out when the fix and its regression test are in; it holds no state and needs no prior setup. ask-espo routes "Something's broken" here.

Two neighbours matter. improve-codebase-architecture takes the handoff when the real finding is that the code has no seam to lock the bug down — the recommendation is made after the fix is in, when there is more information. triage sits upstream of it for bugs that arrive as raw reports from other people, and does a shallower version of the same first two phases.

Try it once

Use the core behavior in one conversation before installation. The repeatable Skill package is the primary path when you want the behavior available across future work.

Copyable text
Diagnose this bug without guessing. First create or identify one red feedback loop for the failure, minimize it, rank hypotheses, instrument the strongest one, then fix and add a regression test. Redact secrets from all evidence.

Source & license

Released in EspoAI Skills v0.1.2; adapted from mattpocock/skills v1.2.3. The released package is skills/engineering/diagnosing-bugs/SKILL.md.

The public package is MIT-licensed and pinned here to the exact release commit. View the released EspoAI source

The adapted baseline preserves the upstream copyright, MIT permission notice, and pinned provenance. View the original pinned source

Skill package files

The full Skill text as copied from content/skill-vault/skills/diagnosing-bugs/. Supporting agent configuration files stay in that folder.

SKILL.md

Source text
---
name: diagnosing-bugs
description: Diagnosis loop for hard bugs and performance regressions. Use when the user says "diagnose"/"debug this", or reports something broken/throwing/failing/slow.
---

# Diagnosing Bugs

A discipline for hard bugs. Skip phases only when explicitly justified.

When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear mental model of the relevant modules, and check ADRs in the area you're touching.

## Redact

This skill has you show commands, outputs and captured artifacts. **Redact every secret first** — write `<REDACTED>` in its place. Build loops against env vars, so the credential stays in the environment rather than in what you show. Captured artifacts carry auth headers: quote only the lines that carry the signal.

If the redacted output is not enough to diagnose the bug, say so and ask the user.

## Phase 1 — Build a feedback loop

**This is the skill.** Everything else is mechanical. If you have a **tight** pass/fail signal for the bug — one that goes red on _this_ bug — you will find the cause; bisection, hypothesis-testing, and instrumentation all just consume it. If you don't have one, no amount of staring at code will save you.

Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to give up.**

### Ways to construct one — try them in roughly this order

1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e.
2. **Curl / HTTP script** against a running dev server.
3. **CLI invocation** with a fixture input, diffing stdout against a known-good snapshot.
4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, asserts on DOM/console/network.
5. **Replay a captured trace.** Save a real network request / payload / event log to disk; replay it through the code path in isolation.
6. **Throwaway harness.** Spin up a minimal subset of the system (one service, mocked deps) that exercises the bug code path with a single function call.
7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 random inputs and look for the failure mode.
8. **Bisection harness.** If the bug appeared between two known states (commit, dataset, version), automate "boot at state X, check, repeat" so you can `git bisect run` it.
9. **Differential loop.** Run the same input through old-version vs new-version (or two configs) and diff outputs.
10. **HITL bash script.** Last resort. If a human must click, drive _them_ with `scripts/hitl-loop.template.sh` so the loop is still structured. Captured output feeds back to you.

Build the right feedback loop, and the bug is 90% fixed.

### Tighten the loop

Treat the loop as a product. Once you have _a_ loop, **tighten** it:

- Can I make it faster? (Cache setup, skip unrelated init, narrow the test scope.)
- Can I make the signal sharper? (Assert on the specific symptom, not "didn't crash".)
- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, freeze network.)

A 30-second flaky loop is barely better than no loop; a 2-second deterministic one is tight — a debugging superpower.

### Non-deterministic bugs

The goal is not a clean repro but a **higher reproduction rate**. Loop the trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A 50%-flake bug is debuggable; 1% is not — keep raising the rate until it's debuggable.

### When you genuinely cannot build a loop

Stop and say so explicitly. List what you tried. Ask the user for: (a) access to whatever environment reproduces it, (b) a redacted captured artifact (HAR file, log dump, core dump, screen recording with timestamps), or (c) permission to add temporary production instrumentation. Do **not** proceed to hypothesise without a loop.

### Completion criterion — a tight loop that goes red

Phase 1 is done when the loop is **tight** and **red-capable**: you can name **one command** — a script path, a test invocation, a curl — that you have **already run at least once** (show the invocation and its output, redacted), and that is:

- [ ] **Red-capable** — it drives the actual bug code path and asserts the **user's exact symptom**, so it can go red on this bug and green once fixed. Not "runs without erroring" — it must be able to _catch this specific bug_.
- [ ] **Deterministic** — same verdict every run (flaky bugs: a pinned, high reproduction rate, per above).
- [ ] **Fast** — seconds, not minutes.
- [ ] **Agent-runnable** — you can run it unattended; a human in the loop only via `scripts/hitl-loop.template.sh`.

If you catch yourself reading code to build a theory before this command exists, **stop — jumping straight to a hypothesis is the exact failure this skill prevents.** No red-capable command, no Phase 2.

## Phase 2 — Reproduce + minimise

Run the loop. Watch it go red — the bug appears.

Confirm:

- [ ] The loop produces the failure mode the **user** described — not a different failure that happens to be nearby. Wrong bug = wrong fix.
- [ ] The failure is reproducible across multiple runs (or, for non-deterministic bugs, reproducible at a high enough rate to debug against).
- [ ] You have captured the exact symptom (error message, wrong output, slow timing) so later phases can verify the fix actually addresses it.

### Minimise

Once it's red, shrink the repro to the **smallest scenario that still goes red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running the loop after each cut — keep only what's load-bearing for the failure.

Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer moving parts left to suspect) and becomes the clean regression test in Phase 5.

Done when **every remaining element is load-bearing** — removing any one of them makes the loop go green.

Do not proceed until you have reproduced **and** minimised.

## Phase 3 — Hypothesise

Generate **3–5 ranked hypotheses** before testing any of them. Single-hypothesis generation anchors on the first plausible idea.

Each hypothesis must be **falsifiable**: state the prediction it makes.

> Format: "If <X> is the cause, then <changing Y> will make the bug disappear / <changing Z> will make it worse."

If you cannot state the prediction, the hypothesis is a vibe — discard or sharpen it.

**Show the ranked list to the user before testing.** They often have domain knowledge that re-ranks instantly ("we just deployed a change to #3"), or know hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't block on it — proceed with your ranking if the user is AFK.

## Phase 4 — Instrument

Each probe must map to a specific prediction from Phase 3. **Change one variable at a time.**

Tool preference:

1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats ten logs.
2. **Targeted logs** at the boundaries that distinguish hypotheses.
3. Never "log everything and grep".

**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at the end becomes a single grep. Untagged logs survive; tagged logs die.

**Perf branch.** For performance regressions, logs are usually wrong. Instead: establish a baseline measurement (timing harness, `performance.now()`, profiler, query plan), then bisect. Measure first, fix second.

## Phase 5 — Fix + regression test

Write the regression test **before the fix** — but only if there is a **correct seam** for it.

A correct seam is one where the test exercises the **real bug pattern** as it occurs at the call site. If the only available seam is too shallow (single-caller test when the bug needs multiple callers, unit test that can't replicate the chain that triggered the bug), a regression test there gives false confidence.

**If no correct seam exists, that itself is the finding.** Note it. The codebase architecture is preventing the bug from being locked down. Flag this for the next phase.

If a correct seam exists:

1. Turn the minimised repro into a failing test at that seam.
2. Watch it fail.
3. Apply the fix.
4. Watch it pass.
5. Re-run the Phase 1 feedback loop against the original (un-minimised) scenario.

## Phase 6 — Cleanup + post-mortem

Required before declaring done:

- [ ] Original repro no longer reproduces (re-run the Phase 1 loop)
- [ ] Regression test passes (or absence of seam is documented)
- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix)
- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug location)
- [ ] The hypothesis that turned out correct is stated in the commit / PR message — so the next debugger learns

**Then ask: what would have prevented this bug?** If the answer involves architectural change (no good test seam, tangled callers, hidden coupling) hand off to the `/improve-codebase-architecture` skill with the specifics. Make the recommendation **after** the fix is in, not before — you have more information now than when you started.

Keep going

One useful next step.

Decide what an incoming issue needs next