Skill Vault · Build and check
Find the next useful code improvement
Survey confusing interfaces and duplicated rules before choosing a refactor.
A codebase your agent can inspect and a place to save its report.
Skill name /improve-codebase-architecture
Your next step
Try it for yourself
You’ll need: A codebase your agent can inspect and a place to save its report.
Paste into your coding agent with the relevant project open.
This is a one-off starter for the approach. Installing the full Skill adds its complete instructions.
Survey this codebase for shallow modules, leaking seams, and duplicated policy. Produce a visual candidate report, recommend the highest-leverage deepening opportunity, and do not refactor yet.
What happens nextA visual report compares candidates and recommends one without starting the refactor.
Use it again
Add the full Skill.
The starter lets you try the approach. Installation adds the complete instructions to your AI coding tool.
Copy the setup instructionsFor Codex or Claude Code on your computer
Your next step
Ask your agent to help you install it
You’ll need: Node.js with npx, Git, and the agent you choose. A project folder where you want the Skill available.
Paste this into Codex or Claude Code with your project open. Your agent will help you review and install the package.
Review files and permissions before accepting an install. Adding a Skill does not run it.
Help me install /improve-codebase-architecture from MattEspo23/skills v0.1.2 in this project.
Review the package instructions and supporting files first. Include these Skills: improve-codebase-architecture, codebase-design, grilling, domain-modeling.
Use the command for the agent I am using:
Add to Codex:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill improve-codebase-architecture --skill codebase-design --skill grilling --skill domain-modeling --agent codex --copy
Add to Claude Code:
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill improve-codebase-architecture --skill codebase-design --skill grilling --skill domain-modeling --agent claude-code --copy
Show me where the files will go before making changes. Preserve existing Skills and customizations. Stop if the release, supporting Skills, or target agent cannot be verified. Do not run the Skill or change external services during setup.
After installation, explain how I can use /improve-codebase-architecture and what access it needs.
What happens nextYour agent should review and add /improve-codebase-architecture, /codebase-design, /grilling, /domain-modeling, then explain how to use /improve-codebase-architecture. Stop if a dependency or version cannot be verified.
Includes supporting Skills: /codebase-design, /grilling, /domain-modeling. Source version: v0.1.2.
Prefer a terminal command?
Add to Codex
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill improve-codebase-architecture --skill codebase-design --skill grilling --skill domain-modeling --agent codex --copy
Add to Claude Code
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill improve-codebase-architecture --skill codebase-design --skill grilling --skill domain-modeling --agent claude-code --copy
Installing copies the instructions into your chosen agent. It does not run the Skill or configure project tools.
Keep the package’s supporting files, LICENSE and NOTICE together with the Skill instructions.
Full notes & source materialThe complete original text, examples and reference details.
These are the complete original notes. Planned videos and services mentioned here may not be available yet; the actions above reflect what you can use on this site now.
Find modules worth deepening and present the strongest refactoring candidates as a visual report.
Watch

Skill Vault video planned · Video planned
Install
Public release v0.1.2 — install the latest source or pin the tested release.
Install latest
npx skills@latest add MattEspo23/skills --skill improve-codebase-architecture
Reproducible install
npx skills@latest add 'MattEspo23/skills#v0.1.2' --skill improve-codebase-architecture
Clean installs are verified for Codex and Claude Code. The installer copies editable files into the selected agent; review every Skill before giving it tool access.
This is the public, editable adapted baseline. Its deeper Espo-specific revision and Skill Vault video are planned; the released source can be installed now.
What it does
improve-codebase-architecture surveys a codebase for deepening opportunities — places where a shallow module (an interface nearly as complex as the thing it hides) could become a deep one — writes them up as a self-contained HTML report, and then grills you through whichever one you pick.
It never changes the code. The whole run produces one HTML file in your OS temp directory and a conversation; the refactor itself happens later, in a separate session, through the normal build flow. That is what makes it a survey rather than a refactoring tool, and it is why the skill is worth running on a codebase you are not ready to touch yet.
Two filters keep the report from becoming generic cleanup advice. Every candidate has to pass the deletion test — would removing this module concentrate complexity behind a smaller interface, or just spread it across callers? Only the "concentrates" cases earn a card. And unless you point it at a specific area, it reads recent commit history first and biases the scan toward paths that are actively changing, on the grounds that a deepening in code nobody touches is a refactor you will never cash in.
When to reach for it
You invoke this by typing /improve-codebase-architecture — the agent will not reach for it on its own.
It sits outside the build loop — it is not a step in the main loop but something you run periodically to queue up more work to improve the codebase. The four situations it gets used in:
| Situation | How it is used |
|---|---|
| Routine upkeep | Run it every few days, or whenever a spare moment appears, to stop structure rotting between features. |
| Before a big build | Point it at the spec: "how can we make this change easy?" This is the most effective prompt for it. |
| Brownfield audit | Run it on a large, unstructured or vibe-coded repo to find out what shape it is actually in. |
| Legacy test work | Use it to find the missing seams first, before writing tests against untestable code. |
Where it is confusable with siblings:
- For designing one module you have already chosen, use codebase-design — that is the bench, this is the survey that finds what to put on it.
- For a whole effort too big to hold in one session, use wayfinder.
- For "this specific thing is broken," use diagnosing-bugs. It hands back here when the real finding is that there is no good seam to lock the bug down.
Prerequisites
None to run it. It reads CONTEXT.md and any ADRs in docs/adr/ if they exist, and speaks in your domain's own nouns when they do — a candidate reads as "deepen the Order intake module," not "refactor the FooBarHandler."
It writes in two places. The report goes to <tmpdir>/architecture-review-<timestamp>.html, outside the repo. During the grilling loop it will add or sharpen terms in CONTEXT.md, creating that file if it does not exist, and offer to record a rejected candidate as an ADR so a future run does not re-suggest it.
Depth, and the report that hunts for it
The skill turns on one idea: depth. A deep module puts a lot of behaviour behind a small, stable interface. A shallow one leaks its implementation through an interface nearly as wide as the code beneath it. The report is a hunt for shallowness — pure functions extracted only for testability while the real bugs live in how they are called (no locality), modules leaking across their seams, a concept you cannot understand without opening five files — and a proposal for the deepening that fixes it.
Each candidate is a card: the files involved, the friction, a plain-English solution, the benefit stated in terms of locality and leverage, a before/after diagram, and a strength badge.
| Badge | What it means for you |
|---|---|
Strong | The deletion test passes clearly and the friction is real. Take these seriously. |
Worth exploring | Plausible deepening, but the payoff depends on where the code is going next. |
Speculative | Surfaced for completeness. Most of these are safe to ignore. |
The report ends with a Top recommendation — the one it would tackle first — and then the skill stops and asks which candidate you want to explore. Nothing has been decided at that point, and no code has moved.
What happens after you pick one
Picking a candidate starts a grilling session over it: constraints, what sits behind the seam, which tests survive, what the deepened interface should look like. The output of that session is a decision, not a diff. From there the normal flow applies — take the decision into to-spec, then to-tickets, then implement.
Common questions
It grilled me for an hour about one idea instead of showing me options. Can I turn that off?
Yes — say so when you invoke it ("don't grill me, just show the report"). This is the loudest complaint the skill has. One user put it bluntly: they liked it as "a convenient way to get a thorough analysis of improvements," and after the grilling loop was added found it "borderline unusable," reporting sessions where it proposed a single solution and then asked "10's or 100's of questions." The design intent is that the report comes first and the grill only starts on a candidate you chose, but weaker models skip straight to interviewing you about the first idea they had. Reports in that thread vary sharply by model, and it is an open issue — the skill does not yet have a documented no-grill mode.
The report opened as unstyled raw HTML with no diagrams. What happened?
The report loads Tailwind and Mermaid from CDNs, so it needs network access when you open it, and it breaks silently when something blocks those scripts. The filed case was a security hook demanding SRI hashes: the agent added them, the CDN served different bytes to the browser than to the curl used to compute the hash, and the browser blocked the script. Offline and locked-down environments hit the same wall. The agent cannot see this, because it never renders the page. The workaround is to ask for inline CSS and hand-built SVG diagrams instead of the CDN scaffold. This is an open issue and a real rough edge.
It gave me twelve candidates. Do I work through them in the same session or start a new one?
One candidate per session. Working through several in one conversation fills the context window with the report, the grilling, the domain-model edits and the code changes all at once. The report only lives in a temp file, so carry the candidate itself rather than the file: pick one, grill it, take the decision into /to-spec, and turn the rest into tickets you can pick up independently later. Put the chosen improvement into a spec rather than going straight to implementation. This is a recurring question with no documented workflow in the skill itself.
How should I prompt it?
With the next thing you are building in mind. Where a big build is coming up, point it at the spec and ask "how can we make this change easy?" An unprompted run scans for hot spots on its own, which is fine for routine upkeep, but naming a direction is what makes the report actionable.
Does it work on a large legacy codebase?
Partly. It is strong on big existing codebases lacking consistent structure, and it is the recommended upkeep mechanism after any one-time structural setup. The honest counterweight: users with genuinely out-of-control projects report it "helped a little but still doesn't seem to cut it," and one developer with an eight-year legacy codebase reported the model going in circles where the same skill produces a clean graph on a tidy repo. There is no dedicated /refactor skill for that case yet. If the codebase has no shared vocabulary at all, grill-with-docs to establish one first tends to make this skill's output much better.
How is this different from /codebase-design?
/codebase-design is a reference, not a session driver. It supplies the vocabulary — module, interface, depth, seam, adapter, leverage, locality — and this skill borrows it. Pointing a fresh agent at /codebase-design as the thing to "do" is a known failure: with no process of its own to follow, the agent invents one, re-explores code and runs for a very long time before asking you anything. Drive with this skill; consume that one.
Will it ever tell me the codebase is fine?
Rarely, and you should know that going in. The skill is built to output findings, so the framing pushes it toward producing candidates rather than concluding that nothing is wrong. The strength badges are the defence — a report where everything is Speculative is the skill telling you it found nothing, in the only way it knows how.
Does it work in Codex or another harness?
Partially. The exploration step names Claude Code's Agent tool with subagent_type=Explore directly, so a harness without that tool may skip the parallel exploration rather than substitute its own. The skill still runs; the scan is just less thorough. A harness-neutral rewrite has been proposed but is not merged.
How do I actually implement deep modules in TypeScript?
There is no good answer shipped with the skill. The recurring request is for a TYPESCRIPT.md giving concrete file and module layouts for the principles, and it does not exist. The skill will tell you where a deepening belongs and what should sit behind the seam; translating that into a package or directory structure is currently on you.
It's working if
- The candidates name your domain's concepts, not invented class names — "the Order intake module," not "the FooBarHandler."
- The candidates cluster in files you have edited recently, not in dormant corners of the repo.
- No code changed during the run. The only new file is the HTML report in your temp directory.
- It stops after the report and asks which candidate you want, rather than continuing on its own.
- Each card explains the payoff as locality or leverage, and says which tests get simpler — not just "this is cleaner."
- Rejecting a candidate for a durable reason gets you an offer to record an ADR, so the next run does not re-suggest it.
Where it fits
improve-codebase-architecture is periodic maintenance — run it every few days, outside any chain, to queue up work rather than to do it. Its neighbours are codebase-design, which owns the depth-and-seam vocabulary every candidate is written in, grilling, which walks the decision tree once you have chosen a candidate, and domain-modeling, which keeps CONTEXT.md and the ADRs current as the decision settles. What it produces is an idea, which re-enters the main build flow at grill-with-docs or to-spec. For which skill fits a situation, ask-espo is the router over the whole set.
Try it once
Use the core behavior in one conversation before installation. The repeatable Skill package is the primary path when you want the behavior available across future work.
Survey this codebase for shallow modules, leaking seams, and duplicated policy. Produce a visual candidate report, recommend the highest-leverage deepening opportunity, and do not refactor yet.
Source & license
Released in EspoAI Skills v0.1.2; adapted from mattpocock/skills v1.2.3. The released package is skills/engineering/improve-codebase-architecture/SKILL.md.
The public package is MIT-licensed and pinned here to the exact release commit. View the released EspoAI source
The adapted baseline preserves the upstream copyright, MIT permission notice, and pinned provenance. View the original pinned source
Skill package files
The full Skill text as copied from content/skill-vault/skills/improve-codebase-architecture/. Supporting agent configuration files stay in that folder.
SKILL.md
---
name: improve-codebase-architecture
description: Scan a codebase for deepening opportunities, present them as a visual HTML report, then grill through whichever one you pick.
---
# Improve Codebase Architecture
Surface architectural friction and propose **deepening opportunities** — refactors that turn shallow modules into deep ones. The aim is testability and AI-navigability.
This command is _informed_ by the project's domain model and built on a shared design vocabulary:
- Run the `/codebase-design` skill for the architecture vocabulary (**module**, **interface**, **depth**, **seam**, **adapter**, **leverage**, **locality**) and its principles (the deletion test, "the interface is the test surface", "one adapter = hypothetical seam, two = real"). Use these terms exactly in every suggestion — don't drift into "component," "service," "API," or "boundary."
- The domain language in `CONTEXT.md` gives names to good seams; ADRs in `docs/adr/` record decisions this command should not re-litigate.
## Process
### 1. Explore
**Scope before you scan — YAGNI.** Deepening a module pays off by making future changes to it easier, so put extra weight on the parts of the codebase that have recently changed. Decide *where* to look before you look:
- If the user named a direction — a module, a subsystem, a pain point — take it, and skip the inference below.
- Otherwise, walk back a good stretch of the commit history (`git log --oneline`) to find the codebase's hot spots — the files and areas that keep coming up — and let those paths pull your attention first. If the changes are scattered with no clear hot spot, widen the net.
Read the project's domain glossary (`CONTEXT.md`) and any ADRs in the area you're touching first.
Then spawn a sub-agent to walk the codebase. Don't follow rigid heuristics — explore organically and note where you experience friction:
- Where does understanding one concept require bouncing between many small modules?
- Where are modules **shallow** — interface nearly as complex as the implementation?
- Where have pure functions been extracted just for testability, but the real bugs hide in how they're called (no **locality**)?
- Where do tightly-coupled modules leak across their seams?
- Which parts of the codebase are untested, or hard to test through their current interface?
Apply the **deletion test** to anything you suspect is shallow: would deleting it concentrate complexity, or just move it? A "yes, concentrates" is the signal you want.
### 2. Present candidates as an HTML report
Write a self-contained HTML file to the OS temp directory so nothing lands in the repo. Resolve the temp dir from `$TMPDIR`, falling back to `/tmp` (or `%TEMP%` on Windows), and write to `<tmpdir>/architecture-review-<timestamp>.html` so each run gets a fresh file. Open it for the user — `xdg-open <path>` on Linux, `open <path>` on macOS, `start <path>` on Windows — and tell them the absolute path.
The report uses **Tailwind via CDN** for layout and styling, and **Mermaid via CDN** for diagrams where a graph/flow/sequence reliably communicates the structure. Mix Mermaid with hand-crafted CSS/SVG visuals — use Mermaid when relationships are graph-shaped (call graphs, dependencies, sequences), and hand-built divs/SVG when you want something more editorial (mass diagrams, cross-sections, collapse animations). Each candidate gets a **before/after visualisation**. Be visual.
For each candidate, render a card with:
- **Files** — which files/modules are involved
- **Problem** — why the current architecture is causing friction
- **Solution** — plain English description of what would change
- **Benefits** — explained in terms of locality and leverage, and how tests would improve
- **Before / After diagram** — side-by-side, custom-drawn, illustrating the shallowness and the deepening
- **Recommendation strength** — one of `Strong`, `Worth exploring`, `Speculative`, rendered as a badge
End the report with a **Top recommendation** section: which candidate you'd tackle first and why.
**Use CONTEXT.md vocabulary for the domain, and the `/codebase-design` vocabulary for the architecture.** If `CONTEXT.md` defines "Order," talk about "the Order intake module" — not "the FooBarHandler," and not "the Order service."
**ADR conflicts**: if a candidate contradicts an existing ADR, only surface it when the friction is real enough to warrant revisiting the ADR. Mark it clearly in the card (e.g. a warning callout: _"contradicts ADR-0007 — but worth reopening because…"_). Don't list every theoretical refactor an ADR forbids.
See [HTML-REPORT.md](HTML-REPORT.md) for the full HTML scaffold, diagram patterns, and styling guidance.
Do NOT propose interfaces yet. After the file is written, ask the user: "Which of these would you like to explore?"
### 3. Grilling loop
Once the user picks a candidate, run the `/grilling` skill to walk the decision tree with them — constraints, dependencies, the shape of the deepened module, what sits behind the seam, what tests survive.
Side effects happen inline as decisions crystallize — run the `/domain-modeling` skill to keep the domain model current as you go:
- **Naming a deepened module after a concept not in `CONTEXT.md`?** Add the term to `CONTEXT.md`. Create the file lazily if it doesn't exist.
- **Sharpening a fuzzy term during the conversation?** Update `CONTEXT.md` right there.
- **User rejects the candidate with a load-bearing reason?** Offer an ADR, framed as: _"Want me to record this as an ADR so future architecture reviews don't re-suggest it?"_ Only offer when the reason would actually be needed by a future explorer to avoid re-suggesting the same thing — skip ephemeral reasons ("not worth it right now") and self-evident ones.
- **Want to explore alternative interfaces for the deepened module?** Run the `/codebase-design` skill and use its design-it-twice parallel sub-agent pattern.
HTML-REPORT.md
# HTML Report Format
The architectural review is rendered as a single self-contained HTML file in the OS temp directory. Tailwind and Mermaid both come from CDNs. Mermaid handles graph-shaped diagrams reliably; hand-built divs and inline SVG handle the more editorial visuals (mass diagrams, cross-sections). Mix the two — don't lean on Mermaid for everything, it'll start to look generic.
## Scaffold
```html
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<title>Architecture review — {{repo name}}</title>
<script src="https://cdn.tailwindcss.com"></script>
<script type="module">
import mermaid from "https://cdn.jsdelivr.net/npm/mermaid@11/dist/mermaid.esm.min.mjs";
mermaid.initialize({ startOnLoad: true, theme: "neutral", securityLevel: "loose" });
</script>
<style>
/* small custom layer for things Tailwind doesn't cover cleanly:
dashed seam lines, hand-drawn-feeling arrow heads, etc. */
.seam { stroke-dasharray: 4 4; }
.leak { stroke: #dc2626; }
.deep { background: linear-gradient(135deg, #0f172a, #1e293b); }
</style>
</head>
<body class="bg-stone-50 text-slate-900 font-sans">
<main class="max-w-5xl mx-auto px-6 py-12 space-y-12">
<header>...</header>
<section id="candidates" class="space-y-10">...</section>
<section id="top-recommendation">...</section>
</main>
</body>
</html>
```
## Header
Repo name, date, and a compact legend: solid box = module, dashed line = seam, red arrow = leakage, thick dark box = deep module. No introduction paragraph — straight into the candidates.
## Candidate card
The diagrams carry the weight. Prose is sparse, plain, and uses the glossary terms (from the `/codebase-design` skill) without ceremony.
Each candidate is one `<article>`:
- **Title** — short, names the deepening (e.g. "Collapse the Order intake pipeline").
- **Badge row** — recommendation strength (`Strong` = emerald, `Worth exploring` = amber, `Speculative` = slate), plus a tag for the dependency category (`in-process`, `local-substitutable`, `ports & adapters`, `mock`).
- **Files** — monospaced list, `font-mono text-sm`.
- **Before / After diagram** — the centrepiece. Two columns, side by side. See patterns below.
- **Problem** — one sentence. What hurts.
- **Solution** — one sentence. What changes.
- **Wins** — bullets, ≤6 words each. e.g. "Tests hit one interface", "Pricing logic stops leaking", "Delete 4 shallow wrappers".
- **ADR callout** (if applicable) — one line in an amber-tinted box.
No paragraphs of explanation. If the diagram needs a paragraph to be understood, redraw the diagram.
## Diagram patterns
Pick the pattern that fits the candidate. Mix them. Don't make every diagram look the same — variety is part of the point.
### Mermaid graph (the workhorse for dependencies / call flow)
Use a Mermaid `flowchart` or `graph` when the point is "X calls Y calls Z, and look at the mess." Wrap it in a Tailwind-styled card so it doesn't feel parachuted in. Style with classDef to colour leakage edges red and the deep module dark. Sequence diagrams work well for "before: 6 round-trips; after: 1."
```html
<div class="rounded-lg border border-slate-200 bg-white p-4">
<pre class="mermaid">
flowchart LR
A[OrderHandler] --> B[OrderValidator]
B --> C[OrderRepo]
C -.leak.-> D[PricingClient]
classDef leak stroke:#dc2626,stroke-width:2px;
class C,D leak
</pre>
</div>
```
### Hand-built boxes-and-arrows (when Mermaid's layout fights you)
Modules as `<div>`s with borders and labels. Arrows as inline SVG `<line>` or `<path>` elements positioned absolutely over a relative container. Reach for this when you want the "after" diagram to feel like one thick-bordered deep module with greyed-out internals — Mermaid won't render that with the right weight.
### Cross-section (good for layered shallowness)
Stack horizontal bands (`h-12 border-l-4`) to show layers a call passes through. Before: 6 thin layers each doing nothing. After: 1 thick band labelled with the consolidated responsibility.
### Mass diagram (good for "interface as wide as implementation")
Two rectangles per module — one for interface surface area, one for implementation. Before: interface rectangle is nearly as tall as the implementation rectangle (shallow). After: interface rectangle is short, implementation rectangle is tall (deep).
### Call-graph collapse
Before: a tree of function calls rendered as nested boxes. After: the same tree collapsed into one box, with the now-internal calls shown faded inside it.
## Style guidance
- Lean editorial, not corporate-dashboard. Generous whitespace. Serif optional for headings (`font-serif` works well with stone/slate).
- Colour sparingly: one accent (emerald or indigo) plus red for leakage and amber for warnings.
- Keep diagrams ~320px tall so before/after sits comfortably side by side without scrolling.
- Use `text-xs uppercase tracking-wider` for module labels inside diagrams — they should read as schematic, not as UI.
- The only scripts are the Tailwind CDN and the Mermaid ESM import. The report is otherwise static — no app code, no interactivity beyond Mermaid's own rendering.
## Top recommendation section
One larger card. Candidate name, one sentence on why, anchor link to its card. That's it.
## Tone
Plain English, concise — but the architectural nouns and verbs come straight from the `/codebase-design` skill. Concision is not an excuse to drift.
**Use exactly:** module, interface, implementation, depth, deep, shallow, seam, adapter, leverage, locality.
**Never substitute:** component, service, unit (for module) · API, signature (for interface) · boundary (for seam) · layer, wrapper (for module, when you mean module).
**Phrasings that fit the style:**
- "Order intake module is shallow — interface nearly matches the implementation."
- "Pricing leaks across the seam."
- "Deepen: one interface, one place to test."
- "Two adapters justify the seam: HTTP in prod, in-memory in tests."
**Wins bullets** name the gain in glossary terms: *"locality: bugs concentrate in one module"*, *"leverage: one interface, N call sites"*, *"interface shrinks; implementation absorbs the wrappers"*. Don't write *"easier to maintain"* or *"cleaner code"* — those terms aren't in the glossary and don't earn their place.
No hedging, no throat-clearing, no "it's worth noting that…". If a sentence could be a bullet, make it a bullet. If a bullet could be cut, cut it. If a term isn't in the `/codebase-design` glossary, reach for one that is before inventing a new one.
Keep going