Skip to content

Skill Vault · Build and check

Build one tested behavior at a time

Use a failing test to guide each small implementation step.

A behavior to build and a project with a runnable test setup.

Skill name /tdd

Your next step

Try it for yourself

You’ll need: A behavior to build and a project with a runnable test setup.

Paste into your coding agent with the relevant project open.

This is a one-off starter for the approach. Installing the full Skill adds its complete instructions.

Ready to copy
Build this behavior test-first: [behavior]. Choose the narrowest stable interface, write one failing behavior test, add only enough implementation to pass, then refactor without changing the behavior.

What happens nextOne behavior test fails, the smallest implementation passes it, then the next slice begins.

Use it again

Add the full Skill.

The starter lets you try the approach. Installation adds the complete instructions to your AI coding tool.

Copy the setup instructionsFor Codex or Claude Code on your computer

Your next step

Ask your agent to help you install it

You’ll need: Node.js with npx, Git, and the agent you choose. A project folder where you want the Skill available.

Paste this into Codex or Claude Code with your project open. Your agent will help you review and install the package.

Review files and permissions before accepting an install. Adding a Skill does not run it.

Ready to copy
Help me install /tdd from MattEspo23/skills v0.1.2 in this project.

Review the package instructions and supporting files first. Include these Skills: tdd, codebase-design.

Use the command for the agent I am using:
Add to Codex:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill tdd --skill codebase-design --agent codex --copy

Add to Claude Code:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill tdd --skill codebase-design --agent claude-code --copy

Show me where the files will go before making changes. Preserve existing Skills and customizations. Stop if the release, supporting Skills, or target agent cannot be verified. Do not run the Skill or change external services during setup.

After installation, explain how I can use /tdd and what access it needs.

What happens nextYour agent should review and add /tdd, /codebase-design, then explain how to use /tdd. Stop if a dependency or version cannot be verified.

Includes supporting Skills: /codebase-design. Source version: v0.1.2.

Prefer a terminal command?

Add to Codex

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill tdd --skill codebase-design --agent codex --copy

Add to Claude Code

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill tdd --skill codebase-design --agent claude-code --copy

Inspect the source on GitHub

Installing copies the instructions into your chosen agent. It does not run the Skill or configure project tools.

Keep the package’s supporting files, LICENSE and NOTICE together with the Skill instructions.

Full notes & source materialThe complete original text, examples and reference details.

These are the complete original notes. Planned videos and services mentioned here may not be available yet; the actions above reflect what you can use on this site now.

Build behavior through a red-green-refactor loop at a stable, meaningful seam.

Watch

An EspoAI reference network connecting durable concepts and evidence.

Skill Vault video planned · Video planned

Install

Public release v0.1.2 — install the latest source or pin the tested release.

Install latest

Command
npx skills@latest add MattEspo23/skills --skill tdd

Reproducible install

Command
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill tdd

Clean installs are verified for Codex and Claude Code. The installer copies editable files into the selected agent; review every Skill before giving it tool access.

This is the public, editable adapted baseline. Its deeper Espo-specific revision and Skill Vault video are planned; the released source can be installed now.

What it does

tdd builds a feature or fixes a bug test-first: one failing test, then just enough code to pass it, then the next behaviour. It carries the standards that make that loop produce tests worth keeping — what a good test is, where tests go, what mocks are for, and the three anti-patterns that quietly ruin a suite.

It writes no test at a seam you have not agreed to first. Before any test exists, it names the public boundaries it intends to test at and stops for your confirmation, because testing effort is finite and this is where you spend it on the critical paths instead of on every edge case. The other thing to know is that tdd is a reference, not a driver. It holds the rules of the loop, and something else (you, or implement) runs the session that applies them.

When to reach for it

Type /tdd, or the agent reaches for it automatically when a task fits — building a feature or fixing a bug test-first, or when you say "red-green-refactor".

Reach for it when there is a concrete behaviour to build, with an input and an observable output, and you want tests that survive a refactor.

Your situationWhere to go
A behaviour with defined inputs and outputs — business logic, a request/response contract, a transformation, validationtdd
The behaviour isn't pinned down yetto-spec, which also agrees the test seams before any code is written
The question is really the shape of the interface, not the testscodebase-design
You have a spec or tickets and want the whole build run for youimplement, which drives tdd per ticket
Config, wiring, glue, type annotations, straight CRUD delegationNothing here fits well — see the open gap below

That last row is a real hole, not a stylistic preference. The skill decides where the seams go; nothing in it decides whether a change is worth the loop at all. Run it on a change with no independent source of truth to assert against and you get a test that restates the implementation — the tautological anti-pattern the skill itself warns about, arrived at from the other direction. It is issue #746 and it is open. Until it closes, that judgement is yours or your CLAUDE.md's.

Prerequisites

codebase-design needs to be installed. tdd used to carry its own deep-module and interface-design notes; in v1.0 those were deleted in favour of the shared skill, and tdd now leans on it for interface-design vocabulary. Nothing else — the skill is stateless and writes no files of its own.

The loop, and the seam it runs at

Three words carry this skill.

Red-green. Write the failing test, then only enough code to pass it. No anticipating the test after next. There is no refactor phase: it was dropped in June 2026 because agents essentially never performed it, and because review and implementation work better as separate sessions. Refactoring belongs to code-review.

Vertical slice. One seam, one test, one minimal implementation, then repeat — the first cycle being a tracer bullet that proves a single path end to end. The opposite is horizontal slicing: all the tests first, then all the code. Bulk tests verify imagined behaviour, they check the shape of things rather than what a user does, and they commit you to a test structure before you understand the implementation.

Pre-agreed seam. A seam is the public boundary you observe behaviour at without reaching inside. The rule is absolute: no test at an unconfirmed seam. In the full chain the seams are agreed earlier, during to-spec — "/tdd is told to only work at pre-agreed test seams, /code-review checks that only agreed-upon test seams were used." Invoked on its own, tdd asks you directly.

The three anti-patterns it is written to prevent:

Anti-patternThe tell
Implementation-coupledThe test breaks when you rename an internal function, though behaviour did not change. Mocked internal collaborators, asserted call counts, database queries used to verify instead of the interface.
TautologicalThe expected value is computed the way the code computes it, so the test passes by construction. Expected values have to come from somewhere else — a known-good literal, a worked example, the spec.
Horizontal slicingA batch of tests landed before any implementation.

Mocks are for system boundaries only — external APIs, time, randomness, sometimes the filesystem or the database. Not your own modules.

Common questions

Why doesn't it refactor? The description says "red-green-refactor".

Because the refactor step was removed and the description was not. The removal was deliberate: agents essentially never did it, and keeping implementation and review in separate sessions works better. Whether the result still counts as TDD by the book matters less than whether the loop produces better code. The mismatch between the trigger phrase and the body is filed as issue #589 and is still open, so "red-green-refactor" continues to work as a phrase that fires the skill. What you get is red → green, and refactoring in code-review.

It asked me to choose a test seam and I had no idea which to pick.

This is the most-reported friction with the skill (issue #607). The prompt lists candidate seams by name only, with nothing about what each one catches or misses, so you are choosing between labels. There is no fix shipped yet. The practical workaround is to ask the agent for the trade-offs before answering — what does the component-level seam miss that the integration seam catches, and how much slower is it. It is also why the chain agrees seams up front in to-spec, where you have the whole feature in view rather than one prompt.

It wrote the implementation before the test, even though the skill says red first.

It happens. One user pushed the model on it and got an unusually honest answer: "I knew the skill said 'one test at a time, watch it fail for the right reason' — I read it. I just defaulted to my normal habit." The skill is written to live with this. No instruction makes an agent comply 100% of the time, and forcing the point harder restricts the agent's creativity for little gain — the loop is worth running even when it is not followed strictly, because the results are still better overall. If strict adherence matters for a particular slice, watch the run rather than trusting the skill to enforce it.

Should it write browser or end-to-end tests first?

Usually not, and the skill will not stop it. A user reported the agent writing a Playwright test first, then burning a long loop re-running it and concluding the test was broken for a feature that did not exist yet. Configure this in your CLAUDE.md. Browser tests are slow enough that the red-green feedback loop stops paying for itself; declare in your repo's CLAUDE.md that they are written after the behaviour works.

Does /tdd replace /implement, or the course's /do-work?

No. /tdd documents the methodology; /implement is a very simple work→feedback→commit loop and is the direct stand-in for /do-work. The course's single /do-work step is now split across /implement, /tdd and /code-review. If you are asking which one to run against a ticket, the answer is almost always /implement.

Where did the deep-modules and interface-design guidance go?

Into codebase-design in v1.0, generalised so several skills share one vocabulary. refactoring.md left at the same time; refactoring is now code-review's job, and that skill carries the Fowler smell baseline.

Does it know about my other tickets?

No. Run against one ticket, it will happily propose work that belongs to a sibling ticket, because it has no view of the rest of the issue graph (issue #129). Matt's position is that this is not tdd's job. Passing the spec alongside the ticket helps; right-sizing the tickets in the first place helps more.

It's working if

  • It stops and names the seams it intends to test at, and waits, before any test file exists.
  • One test appears, goes red, gets just enough code to pass, and only then does the next test appear — not a batch of tests followed by a batch of code.
  • Test names read as capabilities ("user can checkout with valid cart"), not as internals ("checkout calls paymentService.process").
  • Expected values in assertions are literals you can trace to the spec, not values recomputed the way the code computes them.
  • Renaming an internal function breaks nothing in the suite.
  • Mocks appear only at external boundaries — the payment API, the clock — and never around your own modules.

Where it fits

tdd is the engine inside the build step of the main chain, rather than a step of its own:

Copyable text
grill-with-docs → to-spec → to-tickets → implement → code-review

to-spec agrees the test seams up front, implement drives tdd per ticket, and code-review checks afterwards that only the agreed seams were used — and owns the refactoring tdd no longer does. Its other neighbour is codebase-design, the shared source of the seam and deep-module vocabulary tdd speaks. You can also reach for it on its own, whenever there is a concrete behaviour to build and no full spec in play. When you are unsure which skill fits your situation, ask-espo routes you.

Try it once

Use the core behavior in one conversation before installation. The repeatable Skill package is the primary path when you want the behavior available across future work.

Copyable text
Build this behavior test-first: [behavior]. Choose the narrowest stable interface, write one failing behavior test, add only enough implementation to pass, then refactor without changing the behavior.

Source & license

Released in EspoAI Skills v0.1.2; adapted from mattpocock/skills v1.2.3. The released package is skills/engineering/tdd/SKILL.md.

The public package is MIT-licensed and pinned here to the exact release commit. View the released EspoAI source

The adapted baseline preserves the upstream copyright, MIT permission notice, and pinned provenance. View the original pinned source

Skill package files

The full Skill text as copied from content/skill-vault/skills/tdd/. Supporting agent configuration files stay in that folder.

SKILL.md

Source text
---
name: tdd
description: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
---

# Test-Driven Development

TDD is the red → green loop. This skill is the reference that makes that loop produce tests worth keeping: what a good test is, where tests go, the anti-patterns, and the rules of the loop. Every section applies on every cycle — consult them before and during the loop, not after.

When exploring the codebase, read `CONTEXT.md` (if it exists) so test names and interface vocabulary match the project's domain language, and respect ADRs in the area you're touching.

## What a good test is

Tests verify behavior through public interfaces, not implementation details. Code can change entirely; tests shouldn't. A good test reads like a specification — "user can checkout with valid cart" tells you exactly what capability exists — and survives refactors because it doesn't care about internal structure.

See [tests.md](tests.md) for examples and [mocking.md](mocking.md) for mocking guidelines.

## Seams — where tests go

A **seam** is the public boundary you test at: the interface where you observe behavior without reaching inside. Tests live at seams, never against internals.

**Test only at pre-agreed seams.** Before writing any test, write down the seams under test and confirm them with the user. No test is written at an unconfirmed seam. You can't test everything — agreeing the seams up front is how testing effort lands on the critical paths and complex logic instead of every edge case.

Ask: "What's the public interface, and which seams should we test?"

When the shape of that interface is itself in question — how deep the module is, where the seam belongs, what the interface should expose — use the `/codebase-design` skill for the vocabulary. It is the shared source of the module, interface, depth, seam, adapter, leverage and locality terms, and it is a reference to consult, not a session to run.

## Anti-patterns

- **Implementation-coupled** — mocks internal collaborators, tests private methods, or verifies through a side channel (querying the database instead of using the interface). The tell: the test breaks when you refactor but behavior hasn't changed.
- **Tautological** — the assertion recomputes the expected value the way the code does (`expect(add(a, b)).toBe(a + b)`, a snapshot derived by hand the same way, a constant asserted equal to itself), so it passes by construction and can never disagree with the code. Expected values must come from an independent source of truth — a known-good literal, a worked example, the spec.
- **Horizontal slicing** — writing all tests first, then all implementation. Bulk tests verify _imagined_ behavior: you test the _shape_ of things rather than user-facing behavior, the tests go insensitive to real changes, and you commit to test structure before understanding the implementation. Work in **vertical slices** instead — one test → one implementation → repeat, each test a **tracer bullet** that responds to what the last cycle taught you.

## Rules of the loop

- **Red before green.** Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
- **One slice at a time.** One seam, one test, one minimal implementation per cycle.
- **Refactoring is not part of the loop.** It belongs to the review stage (see the `code-review` skill), not the red → green implementation cycle.

mocking.md

Source text
# When to Mock

Mock at **system boundaries** only:

- External APIs (payment, email, etc.)
- Databases (sometimes - prefer test DB)
- Time/randomness
- File system (sometimes)

Don't mock:

- Your own classes/modules
- Internal collaborators
- Anything you control

## Designing for Mockability

At system boundaries, design interfaces that are easy to mock:

**1. Use dependency injection**

Pass external dependencies in rather than creating them internally:

```typescript
// Easy to mock
function processPayment(order, paymentClient) {
  return paymentClient.charge(order.total);
}

// Hard to mock
function processPayment(order) {
  const client = new StripeClient(process.env.STRIPE_KEY);
  return client.charge(order.total);
}
```

**2. Prefer SDK-style interfaces over generic fetchers**

Create specific functions for each external operation instead of one generic function with conditional logic:

```typescript
// GOOD: Each function is independently mockable
const api = {
  getUser: (id) => fetch(`/users/${id}`),
  getOrders: (userId) => fetch(`/users/${userId}/orders`),
  createOrder: (data) => fetch('/orders', { method: 'POST', body: data }),
};

// BAD: Mocking requires conditional logic inside the mock
const api = {
  fetch: (endpoint, options) => fetch(endpoint, options),
};
```

The SDK approach means:
- Each mock returns one specific shape
- No conditional logic in test setup
- Easier to see which endpoints a test exercises
- Type safety per endpoint

tests.md

Source text
# Good and Bad Tests

## Good Tests

**Integration-style**: Test through real interfaces, not mocks of internal parts.

```typescript
// GOOD: Tests observable behavior
test("user can checkout with valid cart", async () => {
  const cart = createCart();
  cart.add(product);
  const result = await checkout(cart, paymentMethod);
  expect(result.status).toBe("confirmed");
});
```

Characteristics:

- Tests behavior users/callers care about
- Uses public API only
- Survives internal refactors
- Describes WHAT, not HOW
- One logical assertion per test

## Bad Tests

**Implementation-detail tests**: Coupled to internal structure.

```typescript
// BAD: Tests implementation details
test("checkout calls paymentService.process", async () => {
  const mockPayment = jest.mock(paymentService);
  await checkout(cart, payment);
  expect(mockPayment.process).toHaveBeenCalledWith(cart.total);
});
```

Red flags:

- Mocking internal collaborators
- Testing private methods
- Asserting on call counts/order
- Test breaks when refactoring without behavior change
- Test name describes HOW not WHAT
- Verifying through external means instead of interface

```typescript
// BAD: Bypasses interface to verify
test("createUser saves to database", async () => {
  await createUser({ name: "Alice" });
  const row = await db.query("SELECT * FROM users WHERE name = ?", ["Alice"]);
  expect(row).toBeDefined();
});

// GOOD: Verifies through interface
test("createUser makes user retrievable", async () => {
  const user = await createUser({ name: "Alice" });
  const retrieved = await getUser(user.id);
  expect(retrieved.name).toBe("Alice");
});
```

**Tautological tests**: Expected value restates the implementation, so the test passes by construction.

```typescript
// BAD: Expected value is recomputed the way the code computes it
test("calculateTotal sums line items", () => {
  const items = [{ price: 10 }, { price: 5 }];
  const expected = items.reduce((sum, i) => sum + i.price, 0);
  expect(calculateTotal(items)).toBe(expected);
});

// GOOD: Expected value is an independent, known literal
test("calculateTotal sums line items", () => {
  expect(calculateTotal([{ price: 10 }, { price: 5 }])).toBe(15);
});
```

Keep going

One useful next step.

Build from a settled plan