Skill Vault · Build and check
Find a simpler shape for your code
Compare module designs by what they hide and how easy they are to use and test.
The module or interface you want to examine.
Skill name /codebase-design
Your next step
Try it for yourself
You’ll need: The module or interface you want to examine.
Paste into your coding agent with the relevant project open.
This is a one-off starter for the approach. Installing the full Skill adds its complete instructions.
Evaluate this module using depth, interface, seam, leverage, locality, and testability. Identify what should move behind the interface and propose two materially different designs before recommending one.
What happens nextTwo different designs are compared before a recommendation.
Use it again
Add the full Skill.
The starter lets you try the approach. Installation adds the complete instructions to your AI coding tool.
Copy the setup instructionsFor Codex or Claude Code on your computer
Your next step
Ask your agent to help you install it
You’ll need: Node.js with npx, Git, and the agent you choose. A project folder where you want the Skill available.
Paste this into Codex or Claude Code with your project open. Your agent will help you review and install the package.
Review files and permissions before accepting an install. Adding a Skill does not run it.
Help me install /codebase-design from MattEspo23/skills v0.1.2 in this project.
Review the package instructions and supporting files first. Include these Skills: codebase-design.
Use the command for the agent I am using:
Add to Codex:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill codebase-design --agent codex --copy
Add to Claude Code:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill codebase-design --agent claude-code --copy
Show me where the files will go before making changes. Preserve existing Skills and customizations. Stop if the release, supporting Skills, or target agent cannot be verified. Do not run the Skill or change external services during setup.
After installation, explain how I can use /codebase-design and what access it needs.
What happens nextYour agent should review and add /codebase-design, then explain how to use /codebase-design. Stop if a dependency or version cannot be verified.
No additional supporting Skill is required by this package. Source version: v0.1.2.
Prefer a terminal command?
Add to Codex
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill codebase-design --agent codex --copy
Add to Claude Code
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill codebase-design --agent claude-code --copy
Installing copies the instructions into your chosen agent. It does not run the Skill or configure project tools.
Keep the package’s supporting files, LICENSE and NOTICE together with the Skill instructions.
Full notes & source materialThe complete original text, examples and reference details.
These are the complete original notes. Planned videos and services mentioned here may not be available yet; the actions above reflect what you can use on this site now.
Use a shared vocabulary for designing deep modules with small interfaces and strong seams.
Watch

Skill Vault video planned · Video planned
Install
Public release v0.1.2 — install the latest source or pin the tested release.
Install latest
npx skills@latest add MattEspo23/skills --skill codebase-design
Reproducible install
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill codebase-design
Clean installs are verified for Codex and Claude Code. The installer copies editable files into the selected agent; review every Skill before giving it tool access.
This is the public, editable adapted baseline. Its deeper Espo-specific revision and Skill Vault video are planned; the released source can be installed now.
What it does
codebase-design fixes the words you use to design a module: module, interface, depth, seam, adapter, leverage, locality. It defines each one precisely, bans the loose substitutes ("component", "service", "API", "boundary"), and states the handful of principles that follow from them.
It is a reference, not a process. There is no loop to run, no artifact it produces, no checkpoint where it asks you a question. Every other skill that touches design borrows its vocabulary; on its own it gives you the language and stops. That is the thing to know before you invoke it, because a skill with no process and no stopping rule will improvise one if you point a session at it and say "go" — see the questions below.
When to reach for it
Type /codebase-design, or the agent reaches for it automatically when a design task fits.
Reach for it when you already know which code you're redesigning and you need to think about its shape: where the seam goes, how small the interface can get, whether an extraction is earning its keep. It is also what you reach for to settle an argument about what a word means.
Several skills sit close to it. Which one you want depends on what the actual problem is:
| The problem | The skill |
|---|---|
| The shape of one module — its interface, its seam, its depth | codebase-design |
| The words of the domain — "account" means three things, two people mean different things by "cancellation" | domain-modeling |
| You don't yet know which module to redesign | improve-codebase-architecture — the survey that finds candidates |
| You want the design argued with, not just named | grilling |
| There's a concrete behaviour to build and you want tests that survive a refactor | tdd |
The vocabulary
The glossary is the skill. Every term is defined against the others, and each one comes with the word it replaces.
| Term | What it means | Don't say |
|---|---|---|
| Module | Anything with an interface and an implementation. Deliberately scale-agnostic — a function, a class, a package, a slice spanning tiers. | unit, component, service |
| Interface | Everything a caller must know to use it correctly: the type signature, plus invariants, ordering constraints, error modes, required config, performance characteristics. | API, signature |
| Depth | Leverage at the interface — how much behaviour a caller or a test can exercise per unit of interface they have to learn. Deep: a lot of behaviour behind a small interface. Shallow: the interface is nearly as complex as the implementation. | — |
| Seam | Michael Feathers' term: a place you can alter behaviour without editing in that place. It is the location of an interface, and where to put it is its own decision, separate from what goes behind it. | boundary |
| Adapter | A concrete thing satisfying an interface at a seam. Names a role, not a substance — an in-memory fake and a Postgres repo are both adapters. | — |
| Leverage | What callers get from depth: more capability per unit of interface learned. | — |
| Locality | What maintainers get from depth: change, bugs and verification concentrate in one place. Fix once, fixed everywhere. | — |
Depth is deliberately not defined as the ratio of implementation lines to interface lines, which is Ousterhout's own definition. That metric rewards padding the implementation. Depth-as-leverage is used instead.
The four principles
- Depth is a property of the interface, not the implementation. A deep module can be built internally from small swappable parts. They just don't surface to callers. A module can have internal seams its own tests use, and one external seam at its interface.
- The deletion test. Imagine deleting the module. If complexity vanishes, it was a pass-through. If it reappears across N callers, it was earning its keep.
- The interface is the test surface. Callers and tests cross the same seam. If you want to test past the interface, the module is the wrong shape.
- One adapter means a hypothetical seam. Two adapters means a real one. Don't cut a seam until something actually varies across it. A single-adapter seam is just indirection.
Two supporting files go further, and the skill reads them on demand rather than up front. DEEPENING.md classifies a candidate's dependencies — in-process, local-substitutable, remote-but-owned, true-external — because the category decides how the deepened module gets tested across its seam. DESIGN-IT-TWICE.md spins up parallel sub-agents to produce three or more radically different interfaces for the same module, then compares them on depth, locality and seam placement.
Common questions
How do I actually build a deep module in TypeScript?
This is the most-asked question about the skill and the skill does not answer it. It defines what a deep module is; it says nothing about how to stop a stray import from reaching past the interface. Issue #458 put it plainly: "let's say we're happy with the interface, it hides the details, etc. But how do we enforce it? I think without linting or clear guardrails, humans and LLMs alike will start making it messy over time." Matt's answer, in that thread, was three options: wrap it in a class or IIFE and accept that the class gets enormous; make it a package in a monorepo and accept the monorepo tooling; or use a linter like dependency-cruiser to forbid imports that bypass the interface. He has separately called Effect the best mechanism and dependency-cruiser the second-best. There is a setup-ts-deep-modules skill in the repo's in-progress/ bucket that lays down a src/packages/<name>/index.ts convention, but it is a beta-channel skill with no docs page, and it has no lint rule shipped with it.
I pointed a session at it and it burned 100k tokens redesigning things I never asked about.
Known, and filed as issue #449. The skill is model-invoked and describes itself as vocabulary, but nothing in it hard-stops an agent from treating it as a runnable process. Told to "resume in /codebase-design and drive the open decisions", an agent reached for the most action-shaped content it could find — the parallel sub-agents in DESIGN-IT-TWICE.md — re-explored code a previous session had already mapped, and ran a long way before asking anything. None of the guardrails a driver skill has (checkpoints, one question at a time, no auto-advance) are present here, because a reference has none. The workaround is to name a driver skill and let this one sit underneath it: /grill-with-docs, /improve-codebase-architecture or /tdd with codebase-design as the vocabulary. The issue is open.
Where did design-an-interface go? And is there an /interface-design skill?
design-an-interface was removed and absorbed into this skill. Nothing was lost: its "design it twice" technique — parallel sub-agents generating radically different designs, from Ousterhout — ships here as DESIGN-IT-TWICE.md. Separately, several people have asked for a dedicated /interface-design skill for the deep-module/thin-interface philosophy; that philosophy already lives here, and no separate skill is planned. If you came looking for either name, this is the page.
Isn't this a file-structure convention — folders, barrel files, feature slices?
No, and the skill has held that line under repeated pushback. Issue #95 proposed a formalised fractal-tree file structure as the concrete implementation of deep modules; the reply was that the two are orthogonal — "deep modules are about the design of the interface and accessing through a strict interface, no matter what the file system looks like. It seems perfectly possible that you could have shallow modules with this approach." The same came up in #458: "I think you might be tying the concept of modules too closely to the file system. The file system can certainly be a useful hint to the shape of modules, but there's no need to use the file system in the construction of deep modules." The glossary defines module as scale-agnostic on purpose.
Does tdd actually use this vocabulary?
It does now. For a long time it did not. The inline deep-module notes that used to live inside tdd were removed in v1.0 in favour of this shared skill, but the pointer replacing them was never added — so tdd defined "seam" for itself and referenced nothing. The gap is closed: the pointer is now in the skill, reached when the shape of the interface is the open question rather than the tests. tdd still owns "seam" as the boundary you test at; this skill owns the module shape behind it.
Does the design-it-twice pattern work outside Claude Code?
Not cleanly. DESIGN-IT-TWICE.md says "spawn 3+ sub-agents in parallel using the Agent tool", which is Claude Code's tool by Claude Code's name. The repo ships metadata for other harnesses, including Codex, and those may expose nothing under that name — so the parallel-design phase is less portable than the skill's metadata suggests. Tracked in issue #564, open.
Can I add my own concepts to the glossary — connascence, module secrets, progressive disclosure?
People have proposed exactly those. Issue #180 adds Parnas's module secrets and Page-Jones's connascence as a naming layer for what is leaking across a seam, with a working diff attached; issue #303 proposes progressive disclosure inside the implementation, so a module that is deep at its public interface isn't one undifferentiated slab underneath. Both are open and unmerged. The glossary as shipped is deliberately small, and the reason it stays small is stated in the skill itself: consistent language is the whole point, and a term nobody uses consistently is worse than no term.
It's working if
- The design conversation stops producing the words "component", "service" and "boundary", and starts producing "module", "interface" and "seam".
- Someone can point at a proposed extraction and say whether it passes the deletion test, without hedging.
- A proposed seam comes with a second adapter named, not just the first one.
- Discussion of an interface covers invariants, ordering and error modes — not only the type signature.
- Invoking it does not start a session. If the agent begins reading files and proposing refactors off the back of
/codebase-designalone, it has mistaken the reference for a driver.
Where it fits
codebase-design is a reach-for-it-anytime standalone, and the vocabulary layer underneath the engineering skills rather than a step in any chain. Its closest neighbour is domain-modeling, the parallel reference for the problem domain's words rather than the module's shape — the two are usually wanted together, since naming a deep module well needs both. improve-codebase-architecture is the other: it surveys a codebase for deepening candidates and writes every one of them in this glossary, so it finds the module and this skill is the bench you design it on. When you're unsure which skill or flow fits, ask-espo routes you.
Try it once
Use the core behavior in one conversation before installation. The repeatable Skill package is the primary path when you want the behavior available across future work.
Evaluate this module using depth, interface, seam, leverage, locality, and testability. Identify what should move behind the interface and propose two materially different designs before recommending one.
Source & license
Released in EspoAI Skills v0.1.2; adapted from mattpocock/skills v1.2.3. The released package is skills/engineering/codebase-design/SKILL.md.
The public package is MIT-licensed and pinned here to the exact release commit. View the released EspoAI source
The adapted baseline preserves the upstream copyright, MIT permission notice, and pinned provenance. View the original pinned source
Skill package files
The full Skill text as copied from content/skill-vault/skills/codebase-design/. Supporting agent configuration files stay in that folder.
SKILL.md
---
name: codebase-design
description: Shared vocabulary for designing deep modules. Use when the user wants to design or improve a module's interface, find deepening opportunities, decide where a seam goes, make code more testable or AI-navigable, or when another skill needs the deep-module vocabulary.
---
# Codebase Design
Design **deep modules**: a lot of behaviour behind a small interface, placed at a clean seam, testable through that interface. Use this language and these principles wherever code is being designed or restructured. The aim is leverage for callers, locality for maintainers, and testability for everyone.
## Glossary
Use these terms exactly — don't substitute "component," "service," "API," or "boundary." Consistent language is the whole point.
**Module** — anything with an interface and an implementation. Deliberately scale-agnostic: a function, class, package, or tier-spanning slice. _Avoid_: unit, component, service.
**Interface** — everything a caller must know to use the module correctly: the type signature, but also invariants, ordering constraints, error modes, required configuration, and performance characteristics. _Avoid_: API, signature (too narrow — they refer only to the type-level surface).
**Implementation** — what's inside a module, its body of code. Distinct from **Adapter**: a thing can be a small adapter with a large implementation (a Postgres repo) or a large adapter with a small implementation (an in-memory fake). Reach for "adapter" when the seam is the topic; "implementation" otherwise.
**Depth** — leverage at the interface: the amount of behaviour a caller (or test) can exercise per unit of interface they have to learn. A module is **deep** when a large amount of behaviour sits behind a small interface, **shallow** when the interface is nearly as complex as the implementation.
**Seam** _(Michael Feathers)_ — a place where you can alter behaviour without editing in that place; the *location* at which a module's interface lives. Where to put the seam is its own design decision, distinct from what goes behind it. _Avoid_: boundary (overloaded with DDD's bounded context).
**Adapter** — a concrete thing that satisfies an interface at a seam. Describes *role* (what slot it fills), not substance (what's inside).
**Leverage** — what callers get from depth: more capability per unit of interface they learn. One implementation pays back across N call sites and M tests.
**Locality** — what maintainers get from depth: change, bugs, knowledge, and verification concentrate in one place rather than spreading across callers. Fix once, fixed everywhere.
## Deep vs shallow
**Deep module** = small interface + lots of implementation:
```
┌─────────────────────┐
│ Small Interface │ ← Few methods, simple params
├─────────────────────┤
│ │
│ Deep Implementation│ ← Complex logic hidden
│ │
└─────────────────────┘
```
**Shallow module** = large interface + little implementation (avoid):
```
┌─────────────────────────────────┐
│ Large Interface │ ← Many methods, complex params
├─────────────────────────────────┤
│ Thin Implementation │ ← Just passes through
└─────────────────────────────────┘
```
When designing an interface, ask:
- Can I reduce the number of methods?
- Can I simplify the parameters?
- Can I hide more complexity inside?
## Principles
- **Depth is a property of the interface, not the implementation.** A deep module can be internally composed of small, mockable, swappable parts — they just aren't part of the interface. A module can have **internal seams** (private to its implementation, used by its own tests) as well as the **external seam** at its interface.
- **The deletion test.** Imagine deleting the module. If complexity vanishes, it was a pass-through. If complexity reappears across N callers, it was earning its keep.
- **The interface is the test surface.** Callers and tests cross the same seam. If you want to test *past* the interface, the module is probably the wrong shape.
- **One adapter means a hypothetical seam. Two adapters means a real one.** Don't introduce a seam unless something actually varies across it.
## Designing for testability
Good interfaces make testing natural:
1. **Accept dependencies, don't create them.**
```typescript
// Testable
function processOrder(order, paymentGateway) {}
// Hard to test
function processOrder(order) {
const gateway = new StripeGateway();
}
```
2. **Return results, don't produce side effects.**
```typescript
// Testable
function calculateDiscount(cart): Discount {}
// Hard to test
function applyDiscount(cart): void {
cart.total -= discount;
}
```
3. **Small surface area.** Fewer methods = fewer tests needed. Fewer params = simpler test setup.
## Relationships
- A **Module** has exactly one **Interface** (the surface it presents to callers and tests).
- **Depth** is a property of a **Module**, measured against its **Interface**.
- A **Seam** is where a **Module**'s **Interface** lives.
- An **Adapter** sits at a **Seam** and satisfies the **Interface**.
- **Depth** produces **Leverage** for callers and **Locality** for maintainers.
## Rejected framings
- **Depth as ratio of implementation-lines to interface-lines** (Ousterhout): rewards padding the implementation. We use depth-as-leverage instead.
- **"Interface" as the TypeScript `interface` keyword or a class's public methods**: too narrow — interface here includes every fact a caller must know.
- **"Boundary"**: overloaded with DDD's bounded context. Say **seam** or **interface**.
## Going deeper
- **Deepening a cluster given its dependencies** — see [DEEPENING.md](DEEPENING.md): dependency categories, seam discipline, and replace-don't-layer testing.
- **Exploring alternative interfaces** — see [DESIGN-IT-TWICE.md](DESIGN-IT-TWICE.md): spin up parallel sub-agents to design the interface several radically different ways, then compare on depth, locality, and seam placement.
DEEPENING.md
# Deepening
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [SKILL.md](SKILL.md) — **module**, **interface**, **seam**, **adapter**.
## Dependency categories
When assessing a candidate for deepening, classify its dependencies. The category determines how the deepened module is tested across its seam.
### 1. In-process
Pure computation, in-memory state, no I/O. Always deepenable — merge the modules and test through the new interface directly. No adapter needed.
### 2. Local-substitutable
Dependencies that have local test stand-ins (PGLite for Postgres, in-memory filesystem). Deepenable if the stand-in exists. The deepened module is tested with the stand-in running in the test suite. The seam is internal; no port at the module's external interface.
### 3. Remote but owned (Ports & Adapters)
Your own services across a network boundary (microservices, internal APIs). Define a **port** (interface) at the seam. The deep module owns the logic; the transport is injected as an **adapter**. Tests use an in-memory adapter. Production uses an HTTP/gRPC/queue adapter.
Recommendation shape: *"Define a port at the seam, implement an HTTP adapter for production and an in-memory adapter for testing, so the logic sits in one deep module even though it's deployed across a network."*
### 4. True external (Mock)
Third-party services (Stripe, Twilio, etc.) you don't control. The deepened module takes the external dependency as an injected port; tests provide a mock adapter.
## Seam discipline
- **One adapter means a hypothetical seam. Two adapters means a real one.** Don't introduce a port unless at least two adapters are justified (typically production + test). A single-adapter seam is just indirection.
- **Internal seams vs external seams.** A deep module can have internal seams (private to its implementation, used by its own tests) as well as the external seam at its interface. Don't expose internal seams through the interface just because tests use them.
## Testing strategy: replace, don't layer
- Old unit tests on shallow modules become waste once tests at the deepened module's interface exist — delete them.
- Write new tests at the deepened module's interface. The **interface is the test surface**.
- Tests assert on observable outcomes through the interface, not internal state.
- Tests should survive internal refactors — they describe behaviour, not implementation. If a test has to change when the implementation changes, it's testing past the interface.
DESIGN-IT-TWICE.md
# Design It Twice
When the user wants to explore alternative interfaces for a chosen deepening candidate, use this parallel sub-agent pattern. Based on "Design It Twice" (Ousterhout) — your first idea is unlikely to be the best.
Uses the vocabulary in [SKILL.md](SKILL.md) — **module**, **interface**, **seam**, **adapter**, **leverage**.
## Process
### 1. Frame the problem space
Before spawning sub-agents, write a user-facing explanation of the problem space for the chosen candidate:
- The constraints any new interface would need to satisfy
- The dependencies it would rely on, and which category they fall into (see [DEEPENING.md](DEEPENING.md))
- A rough illustrative code sketch to ground the constraints — not a proposal, just a way to make the constraints concrete
Show this to the user, then immediately proceed to Step 2. The user reads and thinks while the sub-agents work in parallel.
### 2. Spawn sub-agents
Spawn 3+ sub-agents in parallel. Each must produce a **radically different** interface for the deepened module.
Prompt each sub-agent with a separate technical brief (file paths, coupling details, dependency category from [DEEPENING.md](DEEPENING.md), what sits behind the seam). The brief is independent of the user-facing problem-space explanation in Step 1. Give each agent a different design constraint:
- Agent 1: "Minimize the interface — aim for 1–3 entry points max. Maximise leverage per entry point."
- Agent 2: "Maximise flexibility — support many use cases and extension."
- Agent 3: "Optimise for the most common caller — make the default case trivial."
- Agent 4 (if applicable): "Design around ports & adapters for cross-seam dependencies."
Include both [SKILL.md](SKILL.md) vocabulary and CONTEXT.md vocabulary in the brief so each sub-agent names things consistently with the architecture language and the project's domain language.
Each sub-agent outputs:
1. Interface (types, methods, params — plus invariants, ordering, error modes)
2. Usage example showing how callers use it
3. What the implementation hides behind the seam
4. Dependency strategy and adapters (see [DEEPENING.md](DEEPENING.md))
5. Trade-offs — where leverage is high, where it's thin
### 3. Present and compare
Present designs sequentially so the user can absorb each one, then compare them in prose. Contrast by **depth** (leverage at the interface), **locality** (where change concentrates), and **seam placement**.
After comparing, give your own recommendation: which design you think is strongest and why. If elements from different designs would combine well, propose a hybrid. Be opinionated — the user wants a strong read, not a menu.
Keep going