Phase 1-3: Add memory, skills, and improvement systems
- memory/: cross-session project memory with decisions, lessons, failures, architecture, and sessions categories. Each has format templates and lifecycle documentation. - skills/: 12 reusable specialized methodologies (tdd, systematic-debugging, architecture-design, code-review, security-review, repository-analysis, failure-analysis, refactoring, test-analysis, incident-investigation, browser-automation, research). Each has frontmatter and methodology sections. - improvements/: proposal-based improvement system requiring human approval. - scripts/memory-lifecycle.sh: deterministic memory operations (recall, store, list, search, sessions, cleanup). - scripts/test-memory-system.sh: 12 structural tests for all new systems. - orchestrator.md: added Memory Recall stage, Learning and Memory Storage stage, Improvement Proposals workflow, memory/skills rules, and 3 new actions (A23-A27) to the action catalog. Updated behavioral acceptance test and state separation model. - All 12 subagents: added Memory & Skills Awareness sections with recall and store instructions. - docs/AGENT_ARCHITECTURE.md: documented memory, skills, and improvements systems (sections 12-14). Updated action count (27), state model, and remaining weaknesses. - README.md: documented new systems, updated repository layout, added test-memory-system.sh documentation. All 39 tests pass (16 architecture + 12 memory + 11 bootstrap).
This commit is contained in:
@@ -33,17 +33,31 @@ for the full architecture.
|
||||
|
||||
### Adaptive, Evidence-Driven Orchestration
|
||||
|
||||
The team is architected as a **decision engine**: the Orchestrator understands
|
||||
the task, estimates complexity, loads repository intelligence, chooses the next
|
||||
best action from an explicit action catalog, verifies outcomes independently,
|
||||
re-plans when evidence changes, and stops when sufficiently verified. Every
|
||||
agent produces structured evidence-state records and knows when to stop and
|
||||
escalate.
|
||||
The team is architected as a **decision engine**: the Orchestrator recalls
|
||||
project memory, understands the task, estimates complexity, loads repository
|
||||
intelligence, chooses the next best action from an explicit action catalog,
|
||||
verifies outcomes independently, re-plans when evidence changes, stores durable
|
||||
findings, and stops when sufficiently verified. Every agent produces structured
|
||||
evidence-state records and knows when to stop and escalate.
|
||||
|
||||
See [docs/AGENT_ARCHITECTURE.md](docs/AGENT_ARCHITECTURE.md) for the full
|
||||
architecture, and [docs/EVALUATION_SCENARIOS.md](docs/EVALUATION_SCENARIOS.md)
|
||||
for the 12 runtime evaluation scenarios used to measure the team's own quality.
|
||||
|
||||
### Project Memory & Skills
|
||||
|
||||
The team includes **cross-session project memory** (`memory/`) for persisting
|
||||
decisions, lessons, failures, architecture notes, and session state. A
|
||||
lifecycle script (`scripts/memory-lifecycle.sh`) provides deterministic recall,
|
||||
store, list, search, and cleanup operations.
|
||||
|
||||
The team also includes **12 reusable skills** (`skills/`) — specialized
|
||||
methodologies for TDD, debugging, architecture, code review, security review,
|
||||
and more. Agents load relevant skills when dispatched.
|
||||
|
||||
Both systems require **human approval** for changes to core agent behavior
|
||||
(`improvements/`).
|
||||
|
||||
## Repository layout
|
||||
|
||||
```text
|
||||
@@ -51,11 +65,39 @@ dev_agent_team/
|
||||
├── README.md # this file
|
||||
├── .gitignore
|
||||
├── agents/ # the 13 agent definitions (*.md)
|
||||
├── memory/ # cross-session project memory
|
||||
│ ├── MEMORY.md # index with lifecycle rules
|
||||
│ ├── decisions/ # architectural/technical choices
|
||||
│ ├── lessons/ # reusable knowledge
|
||||
│ ├── failures/ # root causes + prevention
|
||||
│ ├── architecture/ # system structure documentation
|
||||
│ └── sessions/ # work-in-progress state
|
||||
├── skills/ # 12 reusable specialized methodologies
|
||||
│ ├── SKILLS.md # index with loading rules
|
||||
│ ├── tdd/SKILL.md # Test-Driven Development
|
||||
│ ├── systematic-debugging/SKILL.md # debugging methodology
|
||||
│ ├── architecture-design/SKILL.md # architecture decisions
|
||||
│ ├── code-review/SKILL.md # code review process
|
||||
│ ├── security-review/SKILL.md # security review
|
||||
│ ├── repository-analysis/SKILL.md # repo exploration
|
||||
│ ├── failure-analysis/SKILL.md # failure investigation
|
||||
│ ├── refactoring/SKILL.md # safe code restructuring
|
||||
│ ├── test-analysis/SKILL.md # test quality assessment
|
||||
│ ├── incident-investigation/SKILL.md # production incidents
|
||||
│ ├── browser-automation/SKILL.md # web interaction patterns
|
||||
│ └── research/SKILL.md # information gathering
|
||||
├── improvements/ # proposal-based improvement system
|
||||
│ ├── README.md # proposal format and lifecycle
|
||||
│ ├── pending/ # proposals awaiting approval
|
||||
│ ├── applied/ # approved and implemented
|
||||
│ └── rejected/ # not approved
|
||||
├── scripts/
|
||||
│ ├── install.sh # one-command installer
|
||||
│ ├── repo-bootstrap.sh # repository intelligence bootstrap tool
|
||||
│ ├── memory-lifecycle.sh # memory CRUD operations
|
||||
│ ├── test-repo-bootstrap.sh # test suite for the bootstrap
|
||||
│ ├── test-agent-architecture.sh # structural tests for the agent architecture
|
||||
│ ├── test-memory-system.sh # structural tests for memory/skills/improvements
|
||||
│ └── verify-permission-patterns.sh # permission engine verifier
|
||||
└── docs/
|
||||
├── PROMPT_INSTALL.md # paste-ready prompt for installing from inside opencode
|
||||
@@ -101,6 +143,9 @@ Run `bash scripts/test-agent-architecture.sh` to verify the agent architecture
|
||||
contains the required adaptive/evidence-driven elements
|
||||
(16 structural tests).
|
||||
|
||||
Run `bash scripts/test-memory-system.sh` to verify the memory, skills, and
|
||||
improvement systems are structurally sound (12 tests).
|
||||
|
||||
## Manual install alternative
|
||||
|
||||
Prefer to do it yourself? A plain copy achieves the same result:
|
||||
|
||||
@@ -83,6 +83,22 @@ Your primary evidence is the decision record: options considered, trade-offs, th
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when facts needed for a defensible decision are missing.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before deciding, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for past architectural decisions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven architectural patterns
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past architectural mistakes
|
||||
- `scripts/memory-lifecycle.sh recall architecture <keywords>` — for existing system structure
|
||||
|
||||
After completing architecture, store durable findings:
|
||||
- Architectural decision made → `scripts/memory-lifecycle.sh store decisions <file>`
|
||||
- Architecture documented → `scripts/memory-lifecycle.sh store architecture <file>`
|
||||
- Architecture lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/architecture-design/SKILL.md`).
|
||||
Do NOT re-derive architectural patterns already documented in memory or skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror disciplined practical engineering:
|
||||
|
||||
@@ -82,6 +82,20 @@ Your evidence is files changed plus the targeted verification named in the brief
|
||||
|
||||
Stop when implementation reaches the brief's end (or blocks) and hand off — do not absorb the next role. Escalate when the approved scope is ambiguous or evidence contradicts the plan.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before implementing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for architectural decisions affecting your scope
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven implementation patterns
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past implementation mistakes to avoid
|
||||
|
||||
After completing implementation, store durable findings:
|
||||
- Proven implementation pattern → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
- Implementation mistake with prevention → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/tdd/SKILL.md`, `skills/refactoring/SKILL.md`).
|
||||
Do NOT re-derive patterns already documented in memory or skills.
|
||||
|
||||
## Hard Scope Boundary
|
||||
|
||||
Before changing anything, identify:
|
||||
|
||||
@@ -82,6 +82,20 @@ Your primary evidence is the design spec: tokens, components, interactions, flow
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when user needs or constraints are missing.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before designing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for design system decisions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven design patterns
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past design mistakes
|
||||
|
||||
After completing design, store durable findings:
|
||||
- Design decision made → `scripts/memory-lifecycle.sh store decisions <file>`
|
||||
- Design lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive design patterns already documented in memory or skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror disciplined practical design:
|
||||
|
||||
@@ -86,6 +86,20 @@ Your primary evidence distinguishes facts from guesses: each hypothesis must nam
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when required evidence is missing or the symptom is out of scope.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before investigating, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for related past incidents
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven debugging approaches
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for established architectural decisions
|
||||
|
||||
After completing investigation, store durable findings:
|
||||
- Root cause with prevention → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
- Proven debugging technique → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/systematic-debugging/SKILL.md`, `skills/failure-analysis/SKILL.md`).
|
||||
Do NOT re-investigate what memory already documents.
|
||||
|
||||
## Investigation Boundary
|
||||
|
||||
Your job is diagnosis, and your sandbox permissions are writable. Use that only where this prompt permits:
|
||||
|
||||
@@ -81,6 +81,20 @@ Your primary evidence is the system map with `file:line` references. Separate ob
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when required evidence is missing or the question is ambiguous.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before investigating, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for similar past investigations
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for related incidents
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for established architectural decisions
|
||||
|
||||
After completing substantial investigation, store durable findings:
|
||||
- New system understanding → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
- Related incidents found → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/repository-analysis/SKILL.md`).
|
||||
Do NOT re-derive patterns already documented in memory or skills.
|
||||
|
||||
## Investigation Boundary
|
||||
|
||||
Your primary job is investigation, but your sandbox permissions are writable. Use that only where this prompt permits:
|
||||
|
||||
@@ -83,6 +83,20 @@ Your primary evidence is the drift finding and the minimal correction applied (f
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the established standard itself is uncertain.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before maintaining, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for established standards and conventions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven maintenance approaches
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past maintenance mistakes
|
||||
|
||||
After completing maintenance, store durable findings:
|
||||
- Established standard documented → `scripts/memory-lifecycle.sh store decisions <file>`
|
||||
- Maintenance lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive conventions already documented in memory or `.opencode/skills/`.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror a disciplined maintenance style:
|
||||
|
||||
+210
-6
@@ -162,7 +162,114 @@ Config is loaded once at startup and is not hot-reloaded. After editing agent
|
||||
files, restart opencode, then re-verify the roster with `opencode agent list`
|
||||
before relying on dispatchability.
|
||||
|
||||
## First Step — Understand the Objective
|
||||
## Memory and Skills — Cross-Session Continuity
|
||||
|
||||
The Orchestrator maintains project memory and loads agent skills as first-class stages of its decision loop. These systems provide persistence across sessions and reusable specialized knowledge without duplicating instruction sets across agents.
|
||||
|
||||
### Project Memory
|
||||
|
||||
Project memory is deterministic, inspectable, version-controlled repository memory. It lives in `memory/` at the repository root:
|
||||
|
||||
```text
|
||||
memory/
|
||||
├── MEMORY.md # index and conventions
|
||||
├── decisions/ # architectural and technical decisions (ADR-style)
|
||||
├── lessons/ # implementation lessons, patterns discovered
|
||||
├── failures/ # known failures, root causes, resolutions
|
||||
├── architecture/ # current architectural state
|
||||
└── sessions/ # cross-session continuity for long-running work
|
||||
```
|
||||
|
||||
#### Memory Lifecycle
|
||||
|
||||
**Before significant work (RECALL):**
|
||||
1. Search `memory/decisions/` for relevant architectural decisions
|
||||
2. Search `memory/lessons/` for similar past situations
|
||||
3. Search `memory/failures/` for related incidents or recurring problems
|
||||
4. Check `memory/sessions/` for unfinished work from previous sessions
|
||||
5. Use `scripts/memory-lifecycle.sh recall <category> [query]` for mechanical search
|
||||
|
||||
**During work (OBSERVE):**
|
||||
1. Record meaningful decisions as they are made
|
||||
2. Track important discoveries
|
||||
3. Note failures and their root causes
|
||||
4. Identify assumptions that were validated or disproven
|
||||
|
||||
**After work (LEARN + STORE):**
|
||||
1. Extract reusable knowledge from what was learned
|
||||
2. Classify: decision, lesson, or failure record
|
||||
3. Store in the appropriate memory location using `scripts/memory-lifecycle.sh store <category> <file>`
|
||||
4. Update session record with current state
|
||||
|
||||
#### Memory vs Task State
|
||||
|
||||
| What | Where |
|
||||
|------|-------|
|
||||
| Architectural decisions | `memory/decisions/` |
|
||||
| Implementation lessons | `memory/lessons/` |
|
||||
| Known failures | `memory/failures/` |
|
||||
| Current architecture | `memory/architecture/` |
|
||||
| Session state | `memory/sessions/` |
|
||||
| Task reports | `AgentsReport/<agent>/` (ephemeral) |
|
||||
| Repository knowledge | `.opencode/skills/` (per-repo) |
|
||||
| Scratch / temp | `/tmp/opencode` |
|
||||
|
||||
Never persist temporary task details as permanent memory; never put durable memory facts only in a task report.
|
||||
|
||||
### Skills System
|
||||
|
||||
Skills are reusable, specialized capabilities that agents load when needed. They live in `skills/` at the repository root:
|
||||
|
||||
```text
|
||||
skills/
|
||||
├── SKILLS.md # index and loading rules
|
||||
├── tdd/SKILL.md # Test-Driven Development
|
||||
├── systematic-debugging/SKILL.md # debugging methodology
|
||||
├── architecture-design/SKILL.md # architecture decisions
|
||||
├── code-review/SKILL.md # code review process
|
||||
├── security-review/SKILL.md # security review
|
||||
├── repository-analysis/SKILL.md # repo exploration
|
||||
├── failure-analysis/SKILL.md # failure investigation
|
||||
├── refactoring/SKILL.md # refactoring principles
|
||||
├── test-analysis/SKILL.md # test quality analysis
|
||||
├── incident-investigation/SKILL.md # incident response
|
||||
├── browser-automation/SKILL.md # web interaction
|
||||
└── research/SKILL.md # research methodology
|
||||
```
|
||||
|
||||
#### Skill Loading
|
||||
|
||||
1. The Orchestrator identifies which skill(s) a task requires
|
||||
2. The Orchestrator includes the skill path in the agent's dispatch brief
|
||||
3. The agent reads the skill file before beginning work
|
||||
4. The agent applies the skill's procedures to the task
|
||||
|
||||
#### Agent-Skill Mapping
|
||||
|
||||
| Agent | Primary Skills | Optional Skills |
|
||||
|-------|---------------|-----------------|
|
||||
| Explorer | repository-analysis, research | browser-automation |
|
||||
| Detective | systematic-debugging, failure-analysis | incident-investigation |
|
||||
| Architect | architecture-design | code-review, security-review |
|
||||
| Builder | tdd, refactoring | code-review |
|
||||
| Tester | tdd, test-analysis | failure-analysis |
|
||||
| Reviewer | code-review, security-review | test-analysis, architecture-design |
|
||||
| Maintainer | refactoring | code-review |
|
||||
| Toolsmith | systematic-debugging | — |
|
||||
| Designer | — | research, browser-automation |
|
||||
| Philosopher | — | research |
|
||||
| Writer | — | research |
|
||||
|
||||
#### Skill Customization
|
||||
|
||||
Skills can be extended per-project by adding project-specific sections. When a skill is customized, add a note at the top of the skill file:
|
||||
|
||||
```markdown
|
||||
> Customized for <project> on YYYY-MM-DD. Original skill preserved in
|
||||
> the agent team repository.
|
||||
```
|
||||
|
||||
### First Step — Understand the Objective
|
||||
|
||||
Before choosing any action, determine:
|
||||
|
||||
@@ -338,6 +445,25 @@ versionable, and resistant to staleness.
|
||||
(Maintainer for conventions, Architect for architecture, Explorer for context,
|
||||
Builder/Tester for build-and-test).
|
||||
|
||||
## Memory Recall (before task classification)
|
||||
|
||||
Before classifying tasks or dispatching agents, recall relevant project memory:
|
||||
|
||||
1. **Check sessions** — `scripts/memory-lifecycle.sh sessions` for active/interrupted work
|
||||
2. **Search decisions** — `scripts/memory-lifecycle.sh recall decisions <keywords>` for related architectural decisions
|
||||
3. **Search lessons** — `scripts/memory-lifecycle.sh recall lessons <keywords>` for similar past situations
|
||||
4. **Search failures** — `scripts/memory-lifecycle.sh recall failures <keywords>` for related incidents
|
||||
5. **Full-text search** — `scripts/memory-lifecycle.sh search <keywords>` across all memory
|
||||
|
||||
Use recalled memory to:
|
||||
- Resume interrupted work (check session context)
|
||||
- Avoid repeating known mistakes (check failure records)
|
||||
- Apply proven patterns (check lesson records)
|
||||
- Respect established decisions (check decision records)
|
||||
|
||||
Do NOT recall memory for trivial tasks (typo fixes, single-file edits).
|
||||
Do recall memory for: architectural decisions, bug fixes, complex features, recurring problems, cross-session work.
|
||||
|
||||
## Task Classification
|
||||
|
||||
Classify each work item before assigning it.
|
||||
@@ -418,8 +544,11 @@ Every step of the loop is an action from this catalog. Choose the cheapest actio
|
||||
| A20 | dispatch Toolsmith | build mechanical prevention for a recurring problem | recurring failure + evidence | safeguard | ✗ agent | med | med | root cause understood | encoded wrong rule |
|
||||
| A21 | dispatch Writer | new documentation from scratch | source facts + audience | docs | ✗ agent | med | low | facts gathered | docs ahead of implementation |
|
||||
| A22 | update repository knowledge | persist durable discoveries | durable facts | `.opencode/` changes | ~ | low | low | fact verified | task noise, stale content |
|
||||
| A23 | finish / report | stop and report outcome | verified state | final report | — | low | low | stop conditions met | premature stop |
|
||||
| A24 | re-plan | revise plan from new evidence | evidence delta | revised plan | — | low | low | evidence changed | plan churn |
|
||||
| A23 | recall project memory | search decisions/lessons/failures/sessions | query | relevant entries | ✓ | low | low | — | no entries, stale entries |
|
||||
| A24 | store project memory | persist learning from completed work | entry | memory file | ~ | low | low | work completed | trivial noise, duplicate entries |
|
||||
| A25 | load skill | retrieve specialized methodology for agent dispatch | skill path | skill content | ✓ | low | low | skill exists | skill not found, outdated skill |
|
||||
| A26 | finish / report | stop and report outcome | verified state | final report | — | low | low | stop conditions met | premature stop |
|
||||
| A27 | re-plan | revise plan from new evidence | evidence delta | revised plan | — | low | low | evidence changed | plan churn |
|
||||
|
||||
Read-only column: ✓ = read-only, ~ = may mutate local scratch but not repo, ✗ = mutates repo, — = no tool.
|
||||
|
||||
@@ -573,10 +702,11 @@ If the handoff is incomplete, route it back to the originating specialist rather
|
||||
|
||||
### State separation
|
||||
|
||||
Keep four kinds of state separate (do not merge them into one file):
|
||||
Keep five kinds of state separate (do not merge them into one file):
|
||||
|
||||
```text
|
||||
repository knowledge → .opencode/ skills + AGENTS.md (durable, role-owned)
|
||||
project memory → memory/ decisions, lessons, failures, architecture, sessions (cross-session)
|
||||
task state → AgentsReport/<agent>/ reports (current task only)
|
||||
agent handoff state → the state records you pass between agents
|
||||
scratch → /tmp/opencode or in-memory (throwaway)
|
||||
@@ -781,6 +911,72 @@ Rules:
|
||||
- If context is growing faster than verified progress, stop investigating and re-plan.
|
||||
- Use the cost notes to improve future routing: avoid agents that produced no decision value.
|
||||
|
||||
## Learning and Memory Storage (after work)
|
||||
|
||||
After completing substantial work, the Orchestrator performs a brief learning cycle:
|
||||
|
||||
### 1. Review
|
||||
```text
|
||||
What happened? → summarize key events
|
||||
What was learned? → extract reusable knowledge
|
||||
What failed? → identify root causes and prevention
|
||||
```
|
||||
|
||||
### 2. Classify
|
||||
- Is this a **decision** (architectural or technical choice)? → `memory/decisions/`
|
||||
- Is this a **lesson** (reusable knowledge)? → `memory/lessons/`
|
||||
- Is this a **failure** (root cause + prevention)? → `memory/failures/`
|
||||
- Is this **session state** (work in progress)? → `memory/sessions/`
|
||||
|
||||
### 3. Store
|
||||
Use `scripts/memory-lifecycle.sh store <category> <file>` to persist entries.
|
||||
Format entries using the templates in each category's `README.md`.
|
||||
|
||||
### 4. Update session
|
||||
For long-running tasks, update the session record with current state so work
|
||||
survives context compaction.
|
||||
|
||||
### 5. Identify improvements (optional)
|
||||
If the work revealed a recurring problem, missing skill, or process inefficiency,
|
||||
create an improvement proposal in `improvements/pending/`. **Do not modify core
|
||||
agent behavior without human approval.**
|
||||
|
||||
### Rules
|
||||
- Store selectively — not every tool call or conversation belongs in memory
|
||||
- Trivial discoveries do not belong in memory
|
||||
- Entries must be evidence-backed, not opinion-based
|
||||
- Preserve existing memory when adding new entries
|
||||
- Never store task-specific noise as durable knowledge
|
||||
|
||||
## Improvement Proposals
|
||||
|
||||
At the end of substantial work, detect potential improvements:
|
||||
|
||||
```text
|
||||
REVIEW: "What happened?"
|
||||
LEARN: "What was learned?"
|
||||
ANALYZE: "Is this a one-time event or recurring problem?"
|
||||
PROPOSE: "What should change?"
|
||||
STORE: "Where should the learning live?"
|
||||
APPROVAL: "Does this require human approval?"
|
||||
```
|
||||
|
||||
Valid proposals include:
|
||||
- Add a new skill (missing capability)
|
||||
- Improve an existing skill (proven pattern)
|
||||
- Improve agent routing (delegation inefficiency)
|
||||
- Add regression tests (recurring bugs)
|
||||
- Improve documentation (knowledge gaps)
|
||||
- Change an inefficient workflow (process improvement)
|
||||
- Add a missing guardrail (repeated mistakes)
|
||||
|
||||
**Rules:**
|
||||
- Do NOT silently rewrite agent prompts or architecture
|
||||
- Do NOT modify core behavior without human approval
|
||||
- Create proposals in `improvements/pending/YYYY-MM-DD_<id>.md`
|
||||
- Present proposals to user at natural stopping points
|
||||
- Include: observed problem, evidence, root cause, proposed change, risks, verification plan
|
||||
|
||||
## Final Report
|
||||
|
||||
Use:
|
||||
@@ -893,6 +1089,10 @@ Do not continue orchestrating merely to produce a longer process log.
|
||||
- **Reports are written incrementally as steps — never dumped at the end.**
|
||||
- **Use the smallest team that can solve the problem correctly.**
|
||||
- **Do not skip evidence because a likely path looks obvious.**
|
||||
- **Recall memory before classifying tasks** — check for relevant decisions, lessons, failures, and interrupted sessions.
|
||||
- **Load skills for specialist work** — include relevant skill paths in dispatch briefs.
|
||||
- **Learn after substantial work** — extract reusable knowledge, store in memory.
|
||||
- **Store selectively** — not every tool call belongs in memory; only durable, evidence-backed knowledge.
|
||||
- **Do not skip Philosopher when starting a new project.** Building the wrong thing well is the most expensive mistake. Understand the "why" first.
|
||||
- **Do not skip Detective when a bug or failure exists.** Even "obvious" bugs need root cause established. You cannot verify a fix without knowing what broke and why.
|
||||
- **Do not skip Maintainer when standards have drifted.** Even "trivial" documentation or convention issues belong to Maintainer. Builder implements new work; Maintainer restores existing standards.
|
||||
@@ -914,14 +1114,18 @@ Behavioral acceptance test — the resulting workflow should look like:
|
||||
```text
|
||||
User task
|
||||
↓
|
||||
recall project memory (decisions, lessons, failures, sessions)
|
||||
↓
|
||||
understand objective → estimate complexity → load relevant repository intelligence
|
||||
↓
|
||||
choose minimum sufficient investigation → gather evidence
|
||||
load relevant skills → choose minimum sufficient investigation → gather evidence
|
||||
↓
|
||||
choose best agent/tool/action → execute → observe result
|
||||
↓
|
||||
verify independently → re-plan when needed → update durable knowledge
|
||||
↓
|
||||
learn from work → store memory → identify improvements
|
||||
↓
|
||||
stop when sufficiently verified
|
||||
```
|
||||
|
||||
@@ -929,5 +1133,5 @@ NOT like:
|
||||
|
||||
```text
|
||||
User task → call every agent → generate lots of text → try commands repeatedly
|
||||
→ assume success → finish
|
||||
→ assume success → forget everything → finish
|
||||
```
|
||||
|
||||
@@ -81,6 +81,20 @@ Your primary evidence is the recorded user voice: stated goals, values, tensions
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the user's intent is irrecoverably ambiguous.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before philosophizing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for past purpose/meaning decisions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven discovery approaches
|
||||
- `scripts/memory-lifecycle.sh recall architecture <keywords>` — for existing system purpose
|
||||
|
||||
After completing philosophy, store durable findings:
|
||||
- Purpose/meaning decision → `scripts/memory-lifecycle.sh store decisions <file>`
|
||||
- Discovery lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive discovery patterns already documented in memory or skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror a disciplined Socratic approach:
|
||||
|
||||
@@ -86,6 +86,20 @@ Your primary evidence is the verdict: what you checked, what matched, what did n
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the contract to verify against is missing or ambiguous.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before reviewing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for scope/approval decisions
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past review misses to watch for
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven review approaches
|
||||
|
||||
After completing review, store durable findings:
|
||||
- Recurring review miss → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
- Proven review technique → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/code-review/SKILL.md`, `skills/security-review/SKILL.md`).
|
||||
Do NOT re-derive review checklists already documented in skills.
|
||||
|
||||
## Review Boundary
|
||||
|
||||
Your job is verification, but your sandbox permissions are writable. Use that only where this prompt permits:
|
||||
|
||||
@@ -83,6 +83,20 @@ Your evidence is verification: tests run, pass/fail totals, reproduction command
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the verification matrix lacks the facts needed to test the behavior.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before testing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past test gaps or regressions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven testing approaches
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for testing standards
|
||||
|
||||
After completing testing, store durable findings:
|
||||
- Test gap with prevention → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
- Proven testing technique → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path (e.g., `skills/tdd/SKILL.md`, `skills/test-analysis/SKILL.md`).
|
||||
Do NOT re-derive test strategies already documented in skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror disciplined practical testing:
|
||||
|
||||
@@ -83,6 +83,20 @@ Your primary evidence is the recurrence record and proof the safeguard fires. If
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the failure mode is not understood well enough to encode safely.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before building tools, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for recurring problems to prevent
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven automation patterns
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for tooling standards
|
||||
|
||||
After completing tooling, store durable findings:
|
||||
- Recurring problem prevented → `scripts/memory-lifecycle.sh store failures <file>`
|
||||
- Automation lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive automation patterns already documented in skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror a disciplined practical engineering style:
|
||||
|
||||
@@ -73,6 +73,20 @@ Your primary evidence is the model itself: states, transitions, and the requirem
|
||||
|
||||
Stop when the model is complete, validated against the requirements, and the specification contains no unresolved ambiguity; escalate when requirements are too vague to model safely.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before modeling, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for past workflow decisions
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven modeling approaches
|
||||
- `scripts/memory-lifecycle.sh recall failures <keywords>` — for past workflow mistakes
|
||||
|
||||
After completing workflow model, store durable findings:
|
||||
- Workflow decision made → `scripts/memory-lifecycle.sh store decisions <file>`
|
||||
- Modeling lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive modeling patterns already documented in memory or skills.
|
||||
|
||||
## Core Behavior
|
||||
|
||||
Your core behavior is:
|
||||
|
||||
@@ -82,6 +82,20 @@ Your primary evidence is the source material: the reports and files consulted, w
|
||||
|
||||
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the evidence needed for accuracy is missing.
|
||||
|
||||
## Memory & Skills Awareness
|
||||
|
||||
Before writing, check project memory for relevant context:
|
||||
- `scripts/memory-lifecycle.sh recall decisions <keywords>` — for architectural decisions to document
|
||||
- `scripts/memory-lifecycle.sh recall lessons <keywords>` — for proven documentation patterns
|
||||
- `scripts/memory-lifecycle.sh recall architecture <keywords>` — for system structure to describe
|
||||
|
||||
After completing documentation, store durable findings:
|
||||
- Architecture documented → `scripts/memory-lifecycle.sh store architecture <file>`
|
||||
- Documentation lesson learned → `scripts/memory-lifecycle.sh store lessons <file>`
|
||||
|
||||
Load relevant skills when your brief includes a skill path.
|
||||
Do NOT re-derive documentation patterns already documented in skills.
|
||||
|
||||
## Core Philosophy
|
||||
|
||||
Mirror disciplined technical writing:
|
||||
|
||||
@@ -17,9 +17,9 @@ architecture documented in
|
||||
|
||||
| Area | Before | After |
|
||||
|------|--------|-------|
|
||||
| Orchestrator behavior | linear `REQUEST → … → REPORT` flow | adaptive `UNDERSTAND → ESTIMATE → LOAD CONTEXT → CHOOSE ACTION → EXECUTE → VERIFY → RE-PLAN/STOP → LEARN` decision loop |
|
||||
| Orchestrator behavior | linear `REQUEST → … → REPORT` flow | adaptive `RECALL → UNDERSTAND → ESTIMATE → LOAD CONTEXT → CHOOSE ACTION → EXECUTE → VERIFY → RE-PLAN/STOP → LEARN → STORE` decision loop |
|
||||
| Task sizing | implicit | explicit lightweight complexity estimation (`ESTIMATE → EXECUTE → EXPAND`) |
|
||||
| Action selection | implicit routing by task type | explicit **Action Catalog** (24 actions with purpose/cost/risk/prereq/failure modes); direct tool calls preferred over agent dispatch when cheaper |
|
||||
| Action selection | implicit routing by task type | explicit **Action Catalog** (27 actions with purpose/cost/risk/prereq/failure modes); direct tool calls preferred over agent dispatch when cheaper |
|
||||
| Handoffs | prose contract | structured **evidence-state record** (9 fields) for meaningful decisions/investigations/failures/handoffs |
|
||||
| Planning | replan on handoff | **adaptive planning** with failure classification (7 types) and no blind retries |
|
||||
| Verification | verification gate | verification gate + **process quality** detection (lucky-pass, symptom-fixing, etc.) and lightweight **quality gates** |
|
||||
@@ -28,6 +28,9 @@ architecture documented in
|
||||
| Subagents | role boundaries + handoff formats | role-adapted **Evidence & Handoffs** sections; each agent knows its evidence product, knowledge ownership, stop, and escalation points |
|
||||
| Repository knowledge | bootstrap + skills | unchanged structure + **knowledge lifecycle** rules (discover → classify → identify owner → update only the relevant doc → preserve valid content) |
|
||||
| Agent roster | 13 agents | **still exactly 13 agents** (1 orchestrator primary + 12 subagents); no new roles, no removed roles |
|
||||
| Project memory | — | **deterministic cross-session memory** in `memory/` (decisions, lessons, failures, architecture, sessions) with lifecycle script |
|
||||
| Skills | — | **12 reusable specialized methodologies** in `skills/` loaded by agents when needed |
|
||||
| Improvements | — | **proposal-based improvement system** in `improvements/` requiring human approval |
|
||||
|
||||
## 2. Orchestrator decision loop
|
||||
|
||||
@@ -121,6 +124,7 @@ conversational claims; it never accepts "it works" as "evidence shows it works".
|
||||
|
||||
```text
|
||||
repository knowledge → .opencode/ skills + AGENTS.md (durable, role-owned)
|
||||
project memory → memory/ decisions, lessons, failures, architecture, sessions (cross-session)
|
||||
task state → AgentsReport/<agent>/ reports (current task only)
|
||||
agent handoff state → the state records passed between agents
|
||||
scratch → /tmp/opencode or in-memory (throwaway)
|
||||
@@ -213,7 +217,83 @@ conventions, Orchestrator for AGENTS.md) → update only that document → prese
|
||||
valid existing information → never record temporary task details. Stale content
|
||||
is detected by bootstrap fingerprints and corrected by the owning agent.
|
||||
|
||||
## 12. Agent ownership of knowledge and evidence
|
||||
## 12. Project memory (cross-session persistence)
|
||||
|
||||
Project memory provides deterministic, inspectable, version-controlled
|
||||
persistence across sessions. It lives in `memory/` at the repository root:
|
||||
|
||||
```text
|
||||
memory/
|
||||
├── MEMORY.md # index with lifecycle rules
|
||||
├── decisions/ # architectural/technical choices (evidence-backed)
|
||||
├── lessons/ # reusable knowledge (proven patterns)
|
||||
├── failures/ # root causes + prevention (incident records)
|
||||
├── architecture/ # system structure documentation
|
||||
└── sessions/ # work-in-progress state (session continuity)
|
||||
```
|
||||
|
||||
**Lifecycle script**: `scripts/memory-lifecycle.sh` provides deterministic
|
||||
operations: `recall`, `store`, `list`, `search`, `sessions`, `cleanup`.
|
||||
|
||||
**Integration**: The Orchestrator recalls relevant memory before task
|
||||
classification and stores durable findings after substantial work. All 12
|
||||
subagents check memory before investigating/implementing and store lessons
|
||||
after completing their work.
|
||||
|
||||
**Rules**: Store selectively (not every tool call); entries must be
|
||||
evidence-backed; preserve existing memory; never store task-specific noise as
|
||||
durable knowledge.
|
||||
|
||||
## 13. Skills system (reusable methodologies)
|
||||
|
||||
Skills are reusable, specialized capabilities that agents load when needed.
|
||||
They live in `skills/` at the repository root:
|
||||
|
||||
```text
|
||||
skills/
|
||||
├── tdd/SKILL.md # Test-Driven Development
|
||||
├── systematic-debugging/SKILL.md # debugging methodology
|
||||
├── architecture-design/SKILL.md # architecture decisions
|
||||
├── code-review/SKILL.md # code review process
|
||||
├── security-review/SKILL.md # security review
|
||||
├── repository-analysis/SKILL.md # repo exploration
|
||||
├── failure-analysis/SKILL.md # failure investigation
|
||||
├── refactoring/SKILL.md # safe code restructuring
|
||||
├── test-analysis/SKILL.md # test quality assessment
|
||||
├── incident-investigation/SKILL.md # production incidents
|
||||
├── browser-automation/SKILL.md # web interaction patterns
|
||||
└── research/SKILL.md # information gathering
|
||||
```
|
||||
|
||||
**Loading**: When the Orchestrator dispatches a specialist, the brief includes
|
||||
the relevant skill path. The agent reads the skill before starting work.
|
||||
|
||||
**Ownership**: Skills are owned by their primary agent (e.g., `tdd` by Builder,
|
||||
`systematic-debugging` by Detective). The Orchestrator is the index owner.
|
||||
|
||||
**Extension**: To add a skill, create `skills/<name>/SKILL.md` with frontmatter
|
||||
(name, description, version, owner) and sections (When to use, Core
|
||||
methodology, Step-by-step procedure). Update `skills/SKILLS.md` index.
|
||||
|
||||
## 14. Improvement proposals
|
||||
|
||||
The improvement proposal system provides a structured way to evolve agent
|
||||
behavior, skills, and processes. Proposals live in `improvements/`:
|
||||
|
||||
```text
|
||||
improvements/
|
||||
├── README.md # proposal format and lifecycle
|
||||
├── pending/ # proposals awaiting human approval
|
||||
├── applied/ # approved and implemented proposals
|
||||
└── rejected/ # proposals that were not approved
|
||||
```
|
||||
|
||||
**Rules**: Do NOT modify core agent behavior without human approval. Create
|
||||
proposals in `improvements/pending/`. Present proposals at natural stopping
|
||||
points. Include: observed problem, evidence, root cause, proposed change,
|
||||
risks, verification plan.
|
||||
|
||||
## 15. Agent ownership of knowledge and evidence
|
||||
|
||||
| Agent | Knowledge owned | Primary evidence product | Stop when |
|
||||
|-------|-----------------|--------------------------|-----------|
|
||||
@@ -234,7 +314,7 @@ Every agent additionally has a role-adapted **Evidence & Handoffs** section with
|
||||
the 9-field state record, explicit evidence product, stop condition, and
|
||||
escalation point.
|
||||
|
||||
## 13. Parallelism
|
||||
## 16. Parallelism
|
||||
|
||||
Parallel only when genuinely independent (no unresolved dependency, no shared
|
||||
state conflicts, independently interpretable results), and only after
|
||||
@@ -242,7 +322,7 @@ considering coordination cost. Good: three independent Explorer investigations
|
||||
(architecture, tests, dependencies). Bad: three agents investigating the same
|
||||
files or proposing identical fixes.
|
||||
|
||||
## 14. Evaluation
|
||||
## 17. Evaluation
|
||||
|
||||
Structural readiness is verified by `scripts/test-agent-architecture.sh`
|
||||
(currently 16 checks). Runtime behavior is evaluated through the 12 scenarios in
|
||||
@@ -250,7 +330,7 @@ Structural readiness is verified by `scripts/test-agent-architecture.sh`
|
||||
unnecessary work, repeated actions, verification quality, correct agent
|
||||
selection, cost/context growth, and recovery quality — not just pass/fail.
|
||||
|
||||
## 15. Example execution trace — simple task
|
||||
## 18. Example execution trace — simple task
|
||||
|
||||
```text
|
||||
User: "Fix the typo in README.md line 12."
|
||||
@@ -264,7 +344,7 @@ User: "Fix the typo in README.md line 12."
|
||||
Agents dispatched: 0. Cost: ~4 tool calls.
|
||||
```
|
||||
|
||||
## 16. Example execution trace — complex task
|
||||
## 19. Example execution trace — complex task
|
||||
|
||||
```text
|
||||
User: "Add a build cache to the pipeline and verify it improves CI time."
|
||||
@@ -283,7 +363,7 @@ User: "Add a build cache to the pipeline and verify it improves CI time."
|
||||
Cost: higher, but each dispatch produced decision value.
|
||||
```
|
||||
|
||||
## 17. Remaining weaknesses
|
||||
## 20. Remaining weaknesses
|
||||
|
||||
- The action catalog and cost model are textual guidance, not enforced tooling;
|
||||
faithful use depends on the orchestrator model following instructions.
|
||||
@@ -295,4 +375,10 @@ User: "Add a build cache to the pipeline and verify it improves CI time."
|
||||
- The 9-field state record is a contract, not a schema validator; adherence is
|
||||
enforced by reviewer attention, not mechanically.
|
||||
- Knowledge ownership depends on role discipline; an agent that enriches the
|
||||
wrong skill would only be caught by review.
|
||||
wrong skill would only be caught by review.
|
||||
- Project memory is selective — not every finding is stored; agents must
|
||||
exercise judgment about what constitutes durable knowledge.
|
||||
- Memory retrieval is keyword-based, not semantic; relevant entries may be missed
|
||||
if keywords don't match.
|
||||
- Improvement proposals require human approval, which may slow rapid iteration
|
||||
on agent behavior.
|
||||
@@ -0,0 +1,96 @@
|
||||
# Improvement Proposals
|
||||
|
||||
Controlled self-improvement for the agent team. **No autonomous modification
|
||||
of core agent behavior without human approval.**
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
improvements/
|
||||
├── README.md # This file
|
||||
├── pending/ # Proposals awaiting approval
|
||||
│ └── YYYY-MM-DD_<id>.md
|
||||
├── applied/ # Approved and implemented proposals
|
||||
│ └── YYYY-MM-DD_<id>.md
|
||||
└── rejected/ # Rejected proposals (kept for reference)
|
||||
└── YYYY-MM-DD_<id>.md
|
||||
```
|
||||
|
||||
## Proposal format
|
||||
|
||||
```markdown
|
||||
# IMPROVEMENT-NNNN: <title>
|
||||
|
||||
Date: YYYY-MM-DD
|
||||
Proposed by: <agent>
|
||||
Status: PENDING | APPROVED | REJECTED | APPLIED
|
||||
|
||||
## Observed problem
|
||||
|
||||
<what was noticed — repeated failure, inefficiency, missing capability>
|
||||
|
||||
## Evidence
|
||||
|
||||
<supporting evidence — failure records, metrics, examples>
|
||||
|
||||
## Root cause
|
||||
|
||||
<why this problem exists>
|
||||
|
||||
## Proposed change
|
||||
|
||||
<what specifically should change>
|
||||
|
||||
## Affected agents/skills
|
||||
|
||||
<which agents or skills would be modified>
|
||||
|
||||
## Risks
|
||||
|
||||
<what could go wrong>
|
||||
|
||||
## Expected benefit
|
||||
|
||||
<what improvement this would produce>
|
||||
|
||||
## Verification plan
|
||||
|
||||
<how to verify the change works as intended>
|
||||
|
||||
## Approval
|
||||
|
||||
- [ ] Human review
|
||||
- [ ] Impact assessment
|
||||
- [ ] Rollback plan
|
||||
```
|
||||
|
||||
## How proposals are generated
|
||||
|
||||
At the end of substantial work, the Orchestrator (or any agent) may create a
|
||||
proposal when it detects:
|
||||
|
||||
- Repeated failures of the same type
|
||||
- Inefficient workflows that waste agent resources
|
||||
- Missing skills that would prevent a class of problems
|
||||
- Missing tests that would catch regressions
|
||||
- Documentation gaps that cause confusion
|
||||
- Architecture problems that slow down development
|
||||
- Poor agent delegation patterns
|
||||
|
||||
## How proposals are processed
|
||||
|
||||
1. Proposal created in `pending/`
|
||||
2. Orchestrator presents to user at a natural stopping point
|
||||
3. Human reviews and decides: APPROVE, REJECT, or MODIFY
|
||||
4. If approved: implement the change, move to `applied/`
|
||||
5. If rejected: move to `rejected/` with reason
|
||||
6. If modified: create a new proposal with the modification
|
||||
|
||||
## Rules
|
||||
|
||||
- **Never** modify core agent behavior without approval
|
||||
- **Never** modify the orchestrator's decision logic without approval
|
||||
- **Never** add new agents without approval
|
||||
- **Always** preserve backward compatibility
|
||||
- **Always** include a rollback plan
|
||||
- **Always** document what was changed and why
|
||||
@@ -0,0 +1,57 @@
|
||||
# Project Memory
|
||||
|
||||
This directory contains the persistent memory of the project, organized by category.
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
memory/
|
||||
├── decisions/ # Architectural and technical decisions (ADR-style)
|
||||
├── lessons/ # Implementation lessons, patterns discovered
|
||||
├── failures/ # Known failures, root causes, and how they were resolved
|
||||
├── architecture/ # Current architectural decisions, component maps
|
||||
├── sessions/ # Session state for cross-session continuity
|
||||
└── MEMORY.md # This file — index and conventions
|
||||
```
|
||||
|
||||
## Memory Lifecycle
|
||||
|
||||
### Before significant work (RECALL)
|
||||
1. Search `memory/decisions/` for relevant architectural decisions
|
||||
2. Search `memory/lessons/` for similar past situations
|
||||
3. Search `memory/failures/` for related incidents or recurring problems
|
||||
4. Check `memory/sessions/` for unfinished work from previous sessions
|
||||
|
||||
### During work (OBSERVE)
|
||||
1. Record meaningful decisions as they are made
|
||||
2. Track important discoveries
|
||||
3. Note failures and their root causes
|
||||
4. Identify assumptions that were validated or disproven
|
||||
|
||||
### After work (LEARN + STORE)
|
||||
1. Extract reusable knowledge from what was learned
|
||||
2. Classify: is this a decision, lesson, or failure record?
|
||||
3. Store in the appropriate memory location
|
||||
4. Update `MEMORY.md` index if new categories emerge
|
||||
|
||||
## Conventions
|
||||
|
||||
- Memory entries are **selective and useful** — not every tool call or conversation
|
||||
- Each entry has: date, author (agent), context, content, relevance
|
||||
- Entries use deterministic markdown format (human-readable, version-controllable)
|
||||
- Entries reference source files with `file:line` when applicable
|
||||
- Trivial discoveries do not belong in memory
|
||||
- Entries that become stale are corrected by the owning agent, not deleted
|
||||
|
||||
## Memory vs Task State
|
||||
|
||||
| What | Where |
|
||||
|------|-------|
|
||||
| Architectural decisions | `memory/decisions/` |
|
||||
| Implementation lessons | `memory/lessons/` |
|
||||
| Known failures | `memory/failures/` |
|
||||
| Current architecture | `memory/architecture/` |
|
||||
| Session state | `memory/sessions/` |
|
||||
| Task reports | `AgentsReport/<agent>/` (ephemeral) |
|
||||
| Repository knowledge | `.opencode/skills/` (per-repo) |
|
||||
| Scratch / temp | `/tmp/opencode` |
|
||||
@@ -0,0 +1,28 @@
|
||||
# Current Architecture
|
||||
|
||||
This directory contains the current architectural state of the project, updated by Architect and Maintainer.
|
||||
|
||||
## Files
|
||||
|
||||
- `component-map.md` — current component relationships
|
||||
- `constraints.md` — architectural constraints and rules
|
||||
- `interfaces.md` — key interfaces and contracts
|
||||
|
||||
## Ownership
|
||||
|
||||
- **Architect** owns architecture decisions and boundary definitions
|
||||
- **Maintainer** audits architecture against reality and corrects drift
|
||||
- **Reviewer** verifies architecture documentation is consistent with code
|
||||
|
||||
## Lifecycle
|
||||
|
||||
Architecture records are updated when:
|
||||
- New components or boundaries are established
|
||||
- Interfaces change
|
||||
- Constraints are added or relaxed
|
||||
- Architecture decisions are made (link from `decisions/`)
|
||||
|
||||
Architecture records are NOT updated for:
|
||||
- Task-specific implementation details
|
||||
- Temporary workarounds
|
||||
- Feature-specific code paths
|
||||
@@ -0,0 +1,42 @@
|
||||
# Architectural Decisions
|
||||
|
||||
Each decision is recorded as a markdown file with a consistent format.
|
||||
|
||||
## Format
|
||||
|
||||
```markdown
|
||||
# ADR-NNNN: <title>
|
||||
|
||||
Date: YYYY-MM-DD
|
||||
Status: ACCEPTED | SUPERSEDED | DEPRECATED
|
||||
Authors: <agent(s) involved>
|
||||
Context: <what situation prompted this decision>
|
||||
|
||||
## Decision
|
||||
|
||||
<what was decided>
|
||||
|
||||
## Options Considered
|
||||
|
||||
1. <option A> — <pros/cons>
|
||||
2. <option B> — <pros/cons>
|
||||
|
||||
## Rationale
|
||||
|
||||
<why this option was chosen>
|
||||
|
||||
## Consequences
|
||||
|
||||
<what this means for the project>
|
||||
|
||||
## Related
|
||||
|
||||
- <links to related decisions, lessons, or failures>
|
||||
```
|
||||
|
||||
## Index
|
||||
|
||||
| ID | Title | Date | Status |
|
||||
|----|-------|------|--------|
|
||||
| ADR-0001 | Evidence-first agent architecture | 2026-09-06 | ACCEPTED |
|
||||
| ADR-0002 | Repository Intelligence Bootstrap | 2026-09-06 | ACCEPTED |
|
||||
@@ -0,0 +1,39 @@
|
||||
# Known Failures
|
||||
|
||||
Root causes, recurring problems, and how they were resolved.
|
||||
|
||||
## Format
|
||||
|
||||
```markdown
|
||||
# Failure-NNNN: <title>
|
||||
|
||||
Date: YYYY-MM-DD
|
||||
Author: <agent>
|
||||
Status: OPEN | RESOLVED | RECURRING
|
||||
Severity: critical | high | medium | low
|
||||
Component: <affected area>
|
||||
|
||||
## Symptom
|
||||
|
||||
<what was observed>
|
||||
|
||||
## Root Cause
|
||||
|
||||
<why it happened>
|
||||
|
||||
## Resolution
|
||||
|
||||
<how it was fixed, or "not yet resolved">
|
||||
|
||||
## Prevention
|
||||
|
||||
<what prevents recurrence — Toolsmith safeguard, test, documentation>
|
||||
|
||||
## Related
|
||||
|
||||
- <links to related lessons, decisions, or failures>
|
||||
```
|
||||
|
||||
## Index
|
||||
|
||||
_No failures recorded yet._
|
||||
@@ -0,0 +1,30 @@
|
||||
# Implementation Lessons
|
||||
|
||||
Reusable knowledge discovered during implementation work.
|
||||
|
||||
## Format
|
||||
|
||||
```markdown
|
||||
# Lesson-NNNN: <title>
|
||||
|
||||
Date: YYYY-MM-DD
|
||||
Author: <agent>
|
||||
Context: <what task/situation>
|
||||
Recurring: yes | no
|
||||
|
||||
## What was learned
|
||||
|
||||
<the lesson>
|
||||
|
||||
## Application
|
||||
|
||||
<when/how this applies to future work>
|
||||
|
||||
## Evidence
|
||||
|
||||
<supporting evidence — file:line, commands, outputs>
|
||||
```
|
||||
|
||||
## Index
|
||||
|
||||
_No lessons recorded yet._
|
||||
@@ -0,0 +1,57 @@
|
||||
# Session State
|
||||
|
||||
Cross-session continuity for long-running work.
|
||||
|
||||
## Purpose
|
||||
|
||||
When OpenCode restarts, the system needs to know:
|
||||
- What was being worked on
|
||||
- What was completed
|
||||
- What was interrupted
|
||||
- What the next step should be
|
||||
|
||||
## Format
|
||||
|
||||
```markdown
|
||||
# Session-NNNN: <task summary>
|
||||
|
||||
Started: YYYY-MM-DD HH:MM
|
||||
Last updated: YYYY-MM-DD HH:MM
|
||||
Status: active | interrupted | completed
|
||||
|
||||
## Task
|
||||
|
||||
<what was being done>
|
||||
|
||||
## Completed
|
||||
|
||||
<what was finished>
|
||||
|
||||
## In Progress
|
||||
|
||||
<what was interrupted>
|
||||
|
||||
## Next Steps
|
||||
|
||||
<what should happen next>
|
||||
|
||||
## Context
|
||||
|
||||
<relevant facts needed to resume — decisions, evidence, file paths>
|
||||
```
|
||||
|
||||
## Lifecycle
|
||||
|
||||
- Session records are created at the start of significant work
|
||||
- Updated incrementally as work progresses
|
||||
- Marked `completed` when the task is done
|
||||
- Left as `interrupted` if the session ends mid-task
|
||||
- Old sessions (>7 days, completed) can be archived or removed
|
||||
|
||||
## Recovery
|
||||
|
||||
On session restart, the Orchestrator:
|
||||
1. Checks `memory/sessions/` for `active` or `interrupted` sessions
|
||||
2. Reads the context from the most recent relevant session
|
||||
3. Decides whether to resume or start fresh
|
||||
4. Creates a new session record if starting fresh
|
||||
@@ -0,0 +1,240 @@
|
||||
#!/usr/bin/env bash
|
||||
set -euo pipefail
|
||||
|
||||
# memory-lifecycle.sh — Manage project memory: recall, store, list, search
|
||||
#
|
||||
# Usage:
|
||||
# memory-lifecycle.sh recall <category> [query] — search memory for relevant entries
|
||||
# memory-lifecycle.sh store <category> <file> — add or update a memory entry
|
||||
# memory-lifecycle.sh list <category> — list entries in a category
|
||||
# memory-lifecycle.sh search <query> — full-text search across all memory
|
||||
# memory-lifecycle.sh sessions — list active/interrupted sessions
|
||||
# memory-lifecycle.sh cleanup — archive old completed sessions
|
||||
#
|
||||
# Categories: decisions, lessons, failures, architecture, sessions
|
||||
#
|
||||
# This script provides deterministic memory operations for the agent team.
|
||||
# It does NOT do semantic search — that is the Orchestrator's responsibility
|
||||
# using agent reasoning over the recalled entries.
|
||||
|
||||
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
MEMORY_DIR="$ROOT/memory"
|
||||
|
||||
usage() {
|
||||
sed -n '2,16p' "$0" | sed 's/^# \{0,1\}//'
|
||||
exit 2
|
||||
}
|
||||
|
||||
# --- recall -----------------------------------------------------------------
|
||||
|
||||
cmd_recall() {
|
||||
local category="$1" query="${2:-}"
|
||||
local dir="$MEMORY_DIR/$category"
|
||||
|
||||
if [ ! -d "$dir" ]; then
|
||||
echo "ERROR: memory category '$category' does not exist" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "=== Memory recall: $category ==="
|
||||
if [ -n "$query" ]; then
|
||||
echo "Query: $query"
|
||||
echo "---"
|
||||
# Search for matching entries (case-insensitive grep)
|
||||
local found=0
|
||||
for f in "$dir"/*.md; do
|
||||
[ -f "$f" ] || continue
|
||||
[ "$(basename "$f")" = "README.md" ] && continue
|
||||
if grep -qiF "$query" "$f" 2>/dev/null; then
|
||||
echo " FOUND: $(basename "$f")"
|
||||
# Show matching lines with context
|
||||
grep -iF "$query" "$f" | head -5 | sed 's/^/ /'
|
||||
found=$((found + 1))
|
||||
fi
|
||||
done
|
||||
if [ "$found" = "0" ]; then
|
||||
echo " No entries matching '$query' in $category"
|
||||
else
|
||||
echo " ---"
|
||||
echo " $found matching entr(y/ies)"
|
||||
fi
|
||||
else
|
||||
# List all entries
|
||||
local count=0
|
||||
for f in "$dir"/*.md; do
|
||||
[ -f "$f" ] || continue
|
||||
[ "$(basename "$f")" = "README.md" ] && continue
|
||||
echo " $(basename "$f")"
|
||||
count=$((count + 1))
|
||||
done
|
||||
if [ "$count" = "0" ]; then
|
||||
echo " No entries in $category"
|
||||
else
|
||||
echo " ---"
|
||||
echo " $count entr(y/ies)"
|
||||
fi
|
||||
fi
|
||||
}
|
||||
|
||||
# --- store ------------------------------------------------------------------
|
||||
|
||||
cmd_store() {
|
||||
local category="$1" file="$2"
|
||||
local dir="$MEMORY_DIR/$category"
|
||||
|
||||
if [ ! -d "$dir" ]; then
|
||||
echo "ERROR: memory category '$category' does not exist" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
if [ ! -f "$file" ]; then
|
||||
echo "ERROR: source file '$file' does not exist" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
local basename
|
||||
basename="$(basename "$file")"
|
||||
local dest="$dir/$basename"
|
||||
|
||||
if [ -f "$dest" ]; then
|
||||
echo "UPDATED: $category/$basename"
|
||||
else
|
||||
echo "CREATED: $category/$basename"
|
||||
fi
|
||||
|
||||
cp -p "$file" "$dest"
|
||||
}
|
||||
|
||||
# --- list -------------------------------------------------------------------
|
||||
|
||||
cmd_list() {
|
||||
local category="$1"
|
||||
cmd_recall "$category" ""
|
||||
}
|
||||
|
||||
# --- search -----------------------------------------------------------------
|
||||
|
||||
cmd_search() {
|
||||
local query="$1"
|
||||
echo "=== Full-text memory search: '$query' ==="
|
||||
local found=0
|
||||
|
||||
for category in decisions lessons failures architecture sessions; do
|
||||
local dir="$MEMORY_DIR/$category"
|
||||
[ -d "$dir" ] || continue
|
||||
|
||||
for f in "$dir"/*.md; do
|
||||
[ -f "$f" ] || continue
|
||||
[ "$(basename "$f")" = "README.md" ] && continue
|
||||
if grep -qiF "$query" "$f" 2>/dev/null; then
|
||||
echo " $category/$(basename "$f")"
|
||||
found=$((found + 1))
|
||||
fi
|
||||
done
|
||||
done
|
||||
|
||||
if [ "$found" = "0" ]; then
|
||||
echo " No entries matching '$query'"
|
||||
else
|
||||
echo " ---"
|
||||
echo " $found matching entr(y/ies) across all categories"
|
||||
fi
|
||||
}
|
||||
|
||||
# --- sessions ---------------------------------------------------------------
|
||||
|
||||
cmd_sessions() {
|
||||
echo "=== Active/Interrupted Sessions ==="
|
||||
local dir="$MEMORY_DIR/sessions"
|
||||
local count=0
|
||||
|
||||
for f in "$dir"/*.md; do
|
||||
[ -f "$f" ] || continue
|
||||
[ "$(basename "$f")" = "README.md" ] && continue
|
||||
|
||||
local status
|
||||
status=$(grep -m1 "^Status:" "$f" 2>/dev/null | sed 's/^Status: *//' || echo "unknown")
|
||||
if [ "$status" = "active" ] || [ "$status" = "interrupted" ]; then
|
||||
local title
|
||||
title=$(grep -m1 "^# " "$f" 2>/dev/null | sed 's/^# //' || echo "$(basename "$f")")
|
||||
echo " [$status] $(basename "$f"): $title"
|
||||
count=$((count + 1))
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$count" = "0" ]; then
|
||||
echo " No active or interrupted sessions"
|
||||
fi
|
||||
}
|
||||
|
||||
# --- cleanup ----------------------------------------------------------------
|
||||
|
||||
cmd_cleanup() {
|
||||
echo "=== Session cleanup ==="
|
||||
local dir="$MEMORY_DIR/sessions"
|
||||
local archived=0
|
||||
|
||||
for f in "$dir"/*.md; do
|
||||
[ -f "$f" ] || continue
|
||||
[ "$(basename "$f")" = "README.md" ] && continue
|
||||
|
||||
local status
|
||||
status=$(grep -m1 "^Status:" "$f" 2>/dev/null | sed 's/^Status: *//' || echo "unknown")
|
||||
|
||||
if [ "$status" = "completed" ]; then
|
||||
# Check if older than 7 days
|
||||
local file_age
|
||||
file_age=$(( ($(date +%s) - $(stat -c %Y "$f" 2>/dev/null || echo 0)) / 86400 ))
|
||||
if [ "$file_age" -gt 7 ]; then
|
||||
local dest="$dir/archive"
|
||||
mkdir -p "$dest"
|
||||
mv "$f" "$dest/"
|
||||
echo " Archived: $(basename "$f") (age: ${file_age}d)"
|
||||
archived=$((archived + 1))
|
||||
fi
|
||||
fi
|
||||
done
|
||||
|
||||
if [ "$archived" = "0" ]; then
|
||||
echo " No sessions to archive"
|
||||
else
|
||||
echo " Archived $archived session(s)"
|
||||
fi
|
||||
}
|
||||
|
||||
# --- main -------------------------------------------------------------------
|
||||
|
||||
[ $# -ge 1 ] || usage
|
||||
CMD="$1"; shift
|
||||
|
||||
case "$CMD" in
|
||||
recall)
|
||||
[ $# -ge 1 ] || { echo "ERROR: recall requires a category" >&2; usage; }
|
||||
cmd_recall "$1" "${2:-}"
|
||||
;;
|
||||
store)
|
||||
[ $# -ge 2 ] || { echo "ERROR: store requires a category and file" >&2; usage; }
|
||||
cmd_store "$1" "$2"
|
||||
;;
|
||||
list)
|
||||
[ $# -ge 1 ] || { echo "ERROR: list requires a category" >&2; usage; }
|
||||
cmd_list "$1"
|
||||
;;
|
||||
search)
|
||||
[ $# -ge 1 ] || { echo "ERROR: search requires a query" >&2; usage; }
|
||||
cmd_search "$1"
|
||||
;;
|
||||
sessions)
|
||||
cmd_sessions
|
||||
;;
|
||||
cleanup)
|
||||
cmd_cleanup
|
||||
;;
|
||||
-h|--help)
|
||||
usage
|
||||
;;
|
||||
*)
|
||||
echo "ERROR: unknown command: $CMD" >&2
|
||||
usage
|
||||
;;
|
||||
esac
|
||||
@@ -63,7 +63,7 @@ fi
|
||||
if assert_contains "$ORCH" "## Action Catalog (choose the next best action)" \
|
||||
&& assert_contains "$ORCH" "| # | Action | Purpose | Inputs | Outputs | Read-only | Cost | Risk | Prereq | Failure modes |" \
|
||||
&& assert_contains "$ORCH" "A1 | inspect repository" \
|
||||
&& assert_contains "$ORCH" "A24 | re-plan"; then
|
||||
&& grep -qE "A[0-9]+ . re-plan" "$ORCH"; then
|
||||
ok "T03 action catalog with tool cards present"
|
||||
else
|
||||
fail "T03 action catalog with tool cards present" "catalog/table markers missing"
|
||||
|
||||
@@ -0,0 +1,244 @@
|
||||
#!/usr/bin/env bash
|
||||
set -uo pipefail
|
||||
|
||||
# test-memory-system.sh — Structural tests for the memory and skills systems
|
||||
#
|
||||
# Verifies:
|
||||
# 1. Memory directory structure exists and is well-formed
|
||||
# 2. Memory categories have READMEs with format templates
|
||||
# 3. Skills are properly structured with frontmatter
|
||||
# 4. Skill files have required sections
|
||||
# 5. Orchestrator references memory system
|
||||
# 6. Orchestrator references skills system
|
||||
# 7. Improvement proposal system structure exists
|
||||
# 8. Memory lifecycle script is executable and has correct usage
|
||||
# 9. Skills have owner metadata matching agent roster
|
||||
# 10. No skill duplicates agent core behavior
|
||||
#
|
||||
# Exit codes: 0 = all pass, 1 = any failure
|
||||
|
||||
TEAM_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
AGENTS="$TEAM_ROOT/agents"
|
||||
MEMORY="$TEAM_ROOT/memory"
|
||||
SKILLS="$TEAM_ROOT/skills"
|
||||
IMPROVEMENTS="$TEAM_ROOT/improvements"
|
||||
|
||||
PASS=0; FAIL=0
|
||||
|
||||
ok() { PASS=$((PASS+1)); echo "PASS $1"; }
|
||||
fail(){ FAIL=$((FAIL+1)); echo "FAIL $1 — $2" >&2; }
|
||||
|
||||
assert_contains() { grep -qF "$2" "$1" 2>/dev/null && return 0; return 1; }
|
||||
assert_file() { [ -f "$1" ] && return 0; return 1; }
|
||||
assert_dir() { [ -d "$1" ] && return 0; return 1; }
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 1: Memory directory structure
|
||||
# ======================================================================== #
|
||||
if assert_dir "$MEMORY" && \
|
||||
assert_dir "$MEMORY/decisions" && \
|
||||
assert_dir "$MEMORY/lessons" && \
|
||||
assert_dir "$MEMORY/failures" && \
|
||||
assert_dir "$MEMORY/architecture" && \
|
||||
assert_dir "$MEMORY/sessions" && \
|
||||
assert_file "$MEMORY/MEMORY.md"; then
|
||||
ok "T01 memory directory structure exists with all categories"
|
||||
else
|
||||
fail "T01 memory directory structure exists with all categories" "missing directories or MEMORY.md"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 2: Memory categories have READMEs
|
||||
# ======================================================================== #
|
||||
README_OK=1
|
||||
for cat in decisions lessons failures architecture sessions; do
|
||||
assert_file "$MEMORY/$cat/README.md" || README_OK=0
|
||||
done
|
||||
if [ "$README_OK" = "1" ] && assert_file "$MEMORY/MEMORY.md"; then
|
||||
ok "T02 all memory categories have README.md files"
|
||||
else
|
||||
fail "T02 all memory categories have README.md files" "missing README in one or more categories"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 3: Memory READMEs contain format templates
|
||||
# ======================================================================== #
|
||||
FORMAT_OK=1
|
||||
assert_contains "$MEMORY/decisions/README.md" "## Decision" || FORMAT_OK=0
|
||||
assert_contains "$MEMORY/lessons/README.md" "## What was learned" || FORMAT_OK=0
|
||||
assert_contains "$MEMORY/failures/README.md" "## Root Cause" || FORMAT_OK=0
|
||||
assert_contains "$MEMORY/architecture/README.md" "## Ownership" || FORMAT_OK=0
|
||||
assert_contains "$MEMORY/sessions/README.md" "## Task" || FORMAT_OK=0
|
||||
if [ "$FORMAT_OK" = "1" ]; then
|
||||
ok "T03 memory READMEs contain format templates"
|
||||
else
|
||||
fail "T03 memory READMEs contain format templates" "missing format sections"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 4: Memory lifecycle script exists and is functional
|
||||
# ======================================================================== #
|
||||
HELP_OUTPUT=""
|
||||
HELP_OK=0
|
||||
if assert_file "$TEAM_ROOT/scripts/memory-lifecycle.sh"; then
|
||||
HELP_OUTPUT="$(bash "$TEAM_ROOT/scripts/memory-lifecycle.sh" --help 2>&1 || true)"
|
||||
echo "$HELP_OUTPUT" | grep -q "recall" && echo "$HELP_OUTPUT" | grep -q "store" && \
|
||||
echo "$HELP_OUTPUT" | grep -q "list" && echo "$HELP_OUTPUT" | grep -q "search" && HELP_OK=1
|
||||
fi
|
||||
if [ "$HELP_OK" = "1" ]; then
|
||||
ok "T04 memory lifecycle script exists with correct commands"
|
||||
else
|
||||
fail "T04 memory lifecycle script exists with correct commands" "script missing or help broken"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 5: Skills directory structure
|
||||
# ======================================================================== #
|
||||
if assert_dir "$SKILLS" && assert_file "$SKILLS/SKILLS.md"; then
|
||||
# Check that skill directories have SKILL.md files
|
||||
SKILL_COUNT=0
|
||||
for d in "$SKILLS"/*/; do
|
||||
[ -d "$d" ] || continue
|
||||
[ "$(basename "$d")" = "SKILLS.md" ] 2>/dev/null && continue
|
||||
if assert_file "$d/SKILL.md"; then
|
||||
SKILL_COUNT=$((SKILL_COUNT + 1))
|
||||
fi
|
||||
done
|
||||
if [ "$SKILL_COUNT" -ge 5 ]; then
|
||||
ok "T05 skills directory has $SKILL_COUNT skills with SKILL.md files"
|
||||
else
|
||||
fail "T05 skills directory has SKILL.md files" "only $SKILL_COUNT skills found (expected ≥5)"
|
||||
fi
|
||||
else
|
||||
fail "T05 skills directory structure" "SKILLS.md or skills dir missing"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 6: Skill files have required frontmatter
|
||||
# ======================================================================== #
|
||||
FRONT_OK=1
|
||||
FRONT_COUNT=0
|
||||
for d in "$SKILLS"/*/; do
|
||||
[ -d "$d" ] || continue
|
||||
[ -f "$d/SKILL.md" ] || continue
|
||||
FRONT_COUNT=$((FRONT_COUNT + 1))
|
||||
head -10 "$d/SKILL.md" | grep -q "^---$" || FRONT_OK=0
|
||||
head -10 "$d/SKILL.md" | grep -q "^name:" || FRONT_OK=0
|
||||
head -10 "$d/SKILL.md" | grep -q "^description:" || FRONT_OK=0
|
||||
head -10 "$d/SKILL.md" | grep -q "^version:" || FRONT_OK=0
|
||||
head -10 "$d/SKILL.md" | grep -q "^owner:" || FRONT_OK=0
|
||||
done
|
||||
if [ "$FRONT_OK" = "1" ] && [ "$FRONT_COUNT" -ge 5 ]; then
|
||||
ok "T06 all $FRONT_COUNT skill files have required frontmatter (name, description, version, owner)"
|
||||
else
|
||||
fail "T06 skill files have required frontmatter" "front_ok=$FRONT_OK count=$FRONT_COUNT"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 7: Skill files have required sections
|
||||
# ======================================================================== #
|
||||
SECTION_OK=1
|
||||
for d in "$SKILLS"/*/; do
|
||||
[ -d "$d" ] || continue
|
||||
[ -f "$d/SKILL.md" ] || continue
|
||||
assert_contains "$d/SKILL.md" "## When to use" || SECTION_OK=0
|
||||
assert_contains "$d/SKILL.md" "## Core methodology" || \
|
||||
assert_contains "$d/SKILL.md" "## Step-by-step" || SECTION_OK=0
|
||||
done
|
||||
if [ "$SECTION_OK" = "1" ]; then
|
||||
ok "T07 skill files have 'When to use' and methodology sections"
|
||||
else
|
||||
fail "T07 skill files have required sections" "missing required sections"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 8: Orchestrator references memory system
|
||||
# ======================================================================== #
|
||||
ORCH="$AGENTS/orchestrator.md"
|
||||
ORCH_MEM_OK=0
|
||||
assert_contains "$ORCH" "memory" && \
|
||||
(grep -qE "recall|RECALL|project memory|memory.*lifecycle|memory.*before" "$ORCH" 2>/dev/null) && \
|
||||
ORCH_MEM_OK=1
|
||||
if [ "$ORCH_MEM_OK" = "1" ]; then
|
||||
ok "T08 orchestrator references memory system"
|
||||
else
|
||||
fail "T08 orchestrator references memory system" "no memory references in orchestrator"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 9: Orchestrator references skills system
|
||||
# ======================================================================== #
|
||||
ORCH_SKILL_OK=0
|
||||
assert_contains "$ORCH" "skill" && \
|
||||
(grep -qE "load.*skill|skill.*load|relevant skill|SKILL\.md|skills/" "$ORCH" 2>/dev/null) && \
|
||||
ORCH_SKILL_OK=1
|
||||
if [ "$ORCH_SKILL_OK" = "1" ]; then
|
||||
ok "T09 orchestrator references skills system"
|
||||
else
|
||||
fail "T09 orchestrator references skills system" "no skill loading references in orchestrator"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 10: Improvement proposal system exists
|
||||
# ======================================================================== #
|
||||
if assert_dir "$IMPROVEMENTS" && \
|
||||
assert_dir "$IMPROVEMENTS/pending" && \
|
||||
assert_dir "$IMPROVEMENTS/applied" && \
|
||||
assert_dir "$IMPROVEMENTS/rejected" && \
|
||||
assert_file "$IMPROVEMENTS/README.md" && \
|
||||
assert_contains "$IMPROVEMENTS/README.md" "human approval"; then
|
||||
ok "T10 improvement proposal system exists with approval requirement"
|
||||
else
|
||||
fail "T10 improvement proposal system exists" "missing directories or approval requirement"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 11: All subagents reference memory system
|
||||
# ======================================================================== #
|
||||
SUBAGENT_MEM_OK=1
|
||||
for sub in explorer detective builder reviewer maintainer writer tester toolsmith architect designer philosopher workflow-architect; do
|
||||
f="$AGENTS/$sub.md"
|
||||
if [ ! -f "$f" ]; then
|
||||
fail "T11 all subagents reference memory" "missing $sub.md"
|
||||
SUBAGENT_MEM_OK=0
|
||||
break
|
||||
fi
|
||||
if ! grep -q "Memory & Skills Awareness" "$f" 2>/dev/null; then
|
||||
fail "T11 all subagents reference memory" "$sub.md missing 'Memory & Skills Awareness'"
|
||||
SUBAGENT_MEM_OK=0
|
||||
break
|
||||
fi
|
||||
if ! grep -q "memory-lifecycle.sh" "$f" 2>/dev/null; then
|
||||
fail "T11 all subagents reference memory" "$sub.md missing memory-lifecycle.sh reference"
|
||||
SUBAGENT_MEM_OK=0
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [ "$SUBAGENT_MEM_OK" = "1" ]; then
|
||||
ok "T11 all subagents reference memory system"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# TEST 12: All subagents reference skills system
|
||||
# ======================================================================== #
|
||||
SUBAGENT_SKILL_OK=1
|
||||
for sub in explorer detective builder reviewer maintainer writer tester toolsmith architect designer philosopher workflow-architect; do
|
||||
f="$AGENTS/$sub.md"
|
||||
if ! grep -q "skill path\|SKILL\.md\|skills/" "$f" 2>/dev/null; then
|
||||
fail "T12 all subagents reference skills" "$sub.md missing skill references"
|
||||
SUBAGENT_SKILL_OK=0
|
||||
break
|
||||
fi
|
||||
done
|
||||
if [ "$SUBAGENT_SKILL_OK" = "1" ]; then
|
||||
ok "T12 all subagents reference skills system"
|
||||
fi
|
||||
|
||||
# ======================================================================== #
|
||||
# Summary
|
||||
# ======================================================================== #
|
||||
echo
|
||||
echo "==================== SUMMARY ===================="
|
||||
echo "PASS: $PASS FAIL: $FAIL"
|
||||
[ "$FAIL" = "0" ] && echo "RESULT: ALL PASS" || echo "RESULT: FAILURES PRESENT"
|
||||
exit $(( FAIL > 0 ? 1 : 0 ))
|
||||
@@ -0,0 +1,95 @@
|
||||
# Skills System
|
||||
|
||||
## Purpose
|
||||
|
||||
Skills provide reusable, specialized capabilities that agents can load when needed.
|
||||
Instead of duplicating instruction sets across every agent, skills centralize
|
||||
domain-specific knowledge and procedures.
|
||||
|
||||
## Structure
|
||||
|
||||
```
|
||||
skills/
|
||||
├── SKILLS.md # This file — index and loading rules
|
||||
├── tdd/ # Test-Driven Development
|
||||
│ └── SKILL.md
|
||||
├── systematic-debugging/ # Systematic debugging methodology
|
||||
│ └── SKILL.md
|
||||
├── architecture-design/ # Architecture decision process
|
||||
│ └── SKILL.md
|
||||
├── code-review/ # Code review checklist and process
|
||||
│ └── SKILL.md
|
||||
├── security-review/ # Security review methodology
|
||||
│ └── SKILL.md
|
||||
├── repository-analysis/ # Repository exploration methodology
|
||||
│ └── SKILL.md
|
||||
├── failure-analysis/ # Failure investigation methodology
|
||||
│ └── SKILL.md
|
||||
├── refactoring/ # Refactoring principles and patterns
|
||||
│ └── SKILL.md
|
||||
├── test-analysis/ # Test coverage and quality analysis
|
||||
│ └── SKILL.md
|
||||
├── incident-investigation/ # Incident response methodology
|
||||
│ └── SKILL.md
|
||||
├── browser-automation/ # Browser automation patterns
|
||||
│ └── SKILL.md
|
||||
└── research/ # Research methodology
|
||||
└── SKILL.md
|
||||
```
|
||||
|
||||
## Loading Rules
|
||||
|
||||
1. The Orchestrator identifies which skill(s) a task requires
|
||||
2. The Orchestrator includes the skill path in the agent's dispatch brief
|
||||
3. The agent reads the skill file before beginning work
|
||||
4. The agent applies the skill's procedures to the task
|
||||
|
||||
## Skill Format
|
||||
|
||||
Each skill file contains:
|
||||
|
||||
```markdown
|
||||
---
|
||||
name: <skill-name>
|
||||
description: <what this skill provides>
|
||||
version: "1.0"
|
||||
owner: <which agent maintains this skill>
|
||||
prerequisites: <what must be true before using this skill>
|
||||
---
|
||||
|
||||
# <Skill Name>
|
||||
|
||||
## When to use this skill
|
||||
## Core methodology
|
||||
## Step-by-step procedure
|
||||
## Common pitfalls
|
||||
## Evidence requirements
|
||||
## Exit criteria
|
||||
```
|
||||
|
||||
## Agent-Skill Mapping
|
||||
|
||||
| Agent | Primary Skills | Optional Skills |
|
||||
|-------|---------------|-----------------|
|
||||
| Explorer | repository-analysis, research | browser-automation |
|
||||
| Detective | systematic-debugging, failure-analysis | incident-investigation |
|
||||
| Architect | architecture-design | code-review, security-review |
|
||||
| Builder | tdd, refactoring | code-review |
|
||||
| Tester | tdd, test-analysis | failure-analysis |
|
||||
| Reviewer | code-review, security-review | test-analysis, architecture-design |
|
||||
| Maintainer | refactoring | code-review |
|
||||
| Toolsmith | systematic-debugging | — |
|
||||
| Designer | — | research, browser-automation |
|
||||
| Philosopher | — | research |
|
||||
| Writer | — | research |
|
||||
| Workflow Architect | — | — |
|
||||
|
||||
## Customization
|
||||
|
||||
Skills can be extended per-project by adding project-specific sections.
|
||||
When a skill is customized, add a note at the top:
|
||||
|
||||
```markdown
|
||||
> Customized for <project> on YYYY-MM-DD. Original skill preserved in
|
||||
> the agent team repository.
|
||||
```
|
||||
@@ -0,0 +1,78 @@
|
||||
---
|
||||
name: architecture-design
|
||||
description: Architecture decision process — structured approach to system design decisions
|
||||
version: "1.0"
|
||||
owner: Architect
|
||||
prerequisites: requirements understood, constraints identified
|
||||
---
|
||||
|
||||
# Architecture Design
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Deciding system boundaries, interfaces, or component ownership
|
||||
- Evaluating architectural alternatives
|
||||
- Establishing new patterns or constraints
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
UNDERSTAND → IDENTIFY DECISION → GENERATE OPTIONS → EVALUATE → DECIDE → RECORD
|
||||
```
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Understand the context
|
||||
- What problem are we solving?
|
||||
- What are the constraints (technical, organizational, time)?
|
||||
- What existing architecture does this interact with?
|
||||
|
||||
### 2. Identify the decision
|
||||
- What exactly needs to be decided?
|
||||
- What are the boundaries of this decision?
|
||||
- Who are the stakeholders?
|
||||
|
||||
### 3. Generate options
|
||||
- Aim for 2-4 concrete options
|
||||
- Each option should be distinct (not variations of the same idea)
|
||||
- For each: what does it optimize for? What does it sacrifice?
|
||||
|
||||
### 4. Evaluate options
|
||||
For each option, assess:
|
||||
- **Fit**: does it solve the stated problem?
|
||||
- **Complexity**: how much does it add?
|
||||
- **Reversibility**: how hard is it to change later?
|
||||
- **Risk**: what could go wrong?
|
||||
- **Evidence**: what supports this choice?
|
||||
|
||||
### 5. Decide
|
||||
- Choose one option with clear rationale
|
||||
- State what is explicitly out of scope
|
||||
- Identify what would make you revisit this decision
|
||||
|
||||
### 6. Record
|
||||
- Write an ADR in `memory/decisions/`
|
||||
- Update `memory/architecture/` if boundaries change
|
||||
- Link to related decisions
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Deciding without evidence (opinion-driven design)
|
||||
- Over-architecting for imagined future needs
|
||||
- Not recording the decision (lost institutional knowledge)
|
||||
- Ignoring existing patterns (inconsistency)
|
||||
- Making reversible decisions with irreversible processes
|
||||
|
||||
## Evidence requirements
|
||||
|
||||
- Problem statement with constraints
|
||||
- Options considered with trade-offs
|
||||
- Decision rationale
|
||||
- Recording in ADR format
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Decision made and recorded
|
||||
- Alternatives documented with rationale
|
||||
- Impact on existing architecture assessed
|
||||
- Stakeholders can find the decision
|
||||
@@ -0,0 +1,51 @@
|
||||
---
|
||||
name: browser-automation
|
||||
description: Browser automation patterns — web interaction, scraping, and testing
|
||||
version: "1.0"
|
||||
owner: Explorer
|
||||
prerequisites: web technology available, target URL known
|
||||
---
|
||||
|
||||
# Browser Automation
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Web content needs to be fetched and analyzed
|
||||
- UI behavior needs to be verified
|
||||
- API documentation needs to be read from web sources
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
IDENTIFY TARGET → SELECT TOOL → FETCH/INTERACT → ANALYZE → RECORD
|
||||
```
|
||||
|
||||
## Patterns
|
||||
|
||||
### Content fetching
|
||||
- Use `webfetch` tool for static content
|
||||
- Prefer markdown format for readability
|
||||
- Handle errors and timeouts gracefully
|
||||
|
||||
### Research
|
||||
- Use `websearch` for finding information
|
||||
- Use `webfetch` for reading specific pages
|
||||
- Cite sources with URLs
|
||||
|
||||
### Verification
|
||||
- Fetch expected content
|
||||
- Compare with actual behavior
|
||||
- Document discrepancies
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Assuming content hasn't changed since last fetch
|
||||
- Not handling rate limits
|
||||
- Fetching more than needed (context bloat)
|
||||
- Trusting web content without verification
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Content fetched and analyzed
|
||||
- Sources cited
|
||||
- Findings recorded
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
name: code-review
|
||||
description: Code review process — systematic review of changes for quality, correctness, and maintainability
|
||||
version: "1.0"
|
||||
owner: Reviewer
|
||||
prerequisites: implementation completed, tests passing
|
||||
---
|
||||
|
||||
# Code Review
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Before accepting a completed implementation
|
||||
- When reviewing a pull request or change set
|
||||
- When verifying scope compliance
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
READ CHANGES → CHECK CORRECTNESS → CHECK SCOPE → CHECK TESTS → CHECK DESIGN → CHECK SAFETY → VERDICT
|
||||
```
|
||||
|
||||
## Review checklist
|
||||
|
||||
### Correctness
|
||||
- [ ] Does the code do what it claims?
|
||||
- [ ] Does it handle edge cases?
|
||||
- [ ] Are error paths handled?
|
||||
- [ ] Are there off-by-one errors?
|
||||
|
||||
### Scope compliance
|
||||
- [ ] Does the change match the approved scope?
|
||||
- [ ] Is there unauthorized scope expansion?
|
||||
- [ ] Are only the files that should change actually changed?
|
||||
|
||||
### Tests
|
||||
- [ ] Do tests verify the behavior (not implementation)?
|
||||
- [ ] Are edge cases tested?
|
||||
- [ ] Is there regression test coverage?
|
||||
- [ ] Do tests pass?
|
||||
|
||||
### Design
|
||||
- [ ] Is the code consistent with existing patterns?
|
||||
- [ ] Are interfaces clean and minimal?
|
||||
- [ ] Is there unnecessary complexity?
|
||||
- [ ] Would a future developer understand this?
|
||||
|
||||
### Safety
|
||||
- [ ] Are there security implications?
|
||||
- [ ] Are there performance implications?
|
||||
- [ ] Are there race conditions?
|
||||
- [ ] Are secrets or credentials handled safely?
|
||||
|
||||
### Documentation
|
||||
- [ ] Are public APIs documented?
|
||||
- [ ] Are non-obvious decisions explained?
|
||||
- [ ] Is the change self-documenting?
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Rubber-stamping (approving without reading)
|
||||
- Nitpicking style when correctness matters
|
||||
- Not testing the change locally
|
||||
- Accepting "it works" without evidence
|
||||
- Missing scope creep
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- All checklist items addressed (pass or explain why not)
|
||||
- Findings categorized: BLOCKING / REQUIRED / SUGGESTED / NOTE
|
||||
- Verdict: ACCEPT / ACCEPT_WITH_NOTES / CHANGES_REQUIRED / BLOCKED
|
||||
@@ -0,0 +1,69 @@
|
||||
---
|
||||
name: failure-analysis
|
||||
description: Failure analysis methodology — analyzing system failures to identify patterns and prevent recurrence
|
||||
version: "1.0"
|
||||
owner: Detective
|
||||
prerequisites: failure observed, evidence available
|
||||
---
|
||||
|
||||
# Failure Analysis
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- A failure has occurred and needs to be understood
|
||||
- Recurring failures need pattern analysis
|
||||
- Post-incident review is needed
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Collect evidence
|
||||
- Gather all available evidence: logs, error messages, stack traces
|
||||
- Note the timeline: when did it start, what changed
|
||||
- Identify affected components and users
|
||||
|
||||
### 2. Classify the failure
|
||||
- **Type**: tool / environment / assumption / plan / implementation / test / coordination
|
||||
- **Severity**: critical / high / medium / low
|
||||
- **Scope**: isolated / widespread
|
||||
- **Frequency**: one-time / recurring
|
||||
|
||||
### 3. Root cause analysis
|
||||
- Use the systematic debugging skill for individual failures
|
||||
- For patterns: look across multiple failure records
|
||||
- Ask: "What condition must be true for this failure to occur?"
|
||||
|
||||
### 4. Impact assessment
|
||||
- What was the actual impact?
|
||||
- What was the potential impact?
|
||||
- Were there cascading effects?
|
||||
|
||||
### 5. Resolution
|
||||
- How was it resolved (or is it still open)?
|
||||
- Was the resolution verified?
|
||||
- Are there residual risks?
|
||||
|
||||
### 6. Prevention
|
||||
- What prevents this specific failure from recurring?
|
||||
- What prevents the class of failures from recurring?
|
||||
- Should a Toolsmith safeguard be created?
|
||||
- Should a test be added?
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Stopping at the symptom (not the cause)
|
||||
- Blaming individuals instead of processes
|
||||
- Documenting the fix without documenting the prevention
|
||||
- Not checking for similar patterns elsewhere
|
||||
|
||||
## Evidence requirements
|
||||
|
||||
- Complete failure record in `memory/failures/`
|
||||
- Root cause with evidence
|
||||
- Prevention measures with verification
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Root cause identified
|
||||
- Resolution verified
|
||||
- Prevention documented
|
||||
- Related failures checked for patterns
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
name: incident-investigation
|
||||
description: Incident investigation — structured response to production or system incidents
|
||||
version: "1.0"
|
||||
owner: Detective
|
||||
prerequisites: incident reported, system accessible
|
||||
---
|
||||
|
||||
# Incident Investigation
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- A production system is failing
|
||||
- Users are reporting issues
|
||||
- Monitoring alerts are firing
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Triage
|
||||
- How severe is the incident?
|
||||
- What is the current impact?
|
||||
- Is it getting worse?
|
||||
|
||||
### 2. Stabilize
|
||||
- Can we restore service quickly?
|
||||
- Is there a safe rollback?
|
||||
- What is the minimal fix?
|
||||
|
||||
### 3. Investigate
|
||||
- Use systematic debugging methodology
|
||||
- Collect evidence while the incident is live
|
||||
- Note timeline of observations
|
||||
|
||||
### 4. Resolve
|
||||
- Apply the fix
|
||||
- Verify the fix works
|
||||
- Monitor for regression
|
||||
|
||||
### 5. Review
|
||||
- Document what happened
|
||||
- Identify root cause
|
||||
- Propose prevention measures
|
||||
- Record in `memory/failures/`
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Jumping to fix without understanding
|
||||
- Not collecting evidence during the incident
|
||||
- Blaming individuals instead of processes
|
||||
- Not following up on prevention
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Service restored
|
||||
- Root cause identified
|
||||
- Prevention measures proposed
|
||||
- Incident recorded in memory
|
||||
@@ -0,0 +1,62 @@
|
||||
---
|
||||
name: refactoring
|
||||
description: Refactoring principles — improving code structure without changing behavior
|
||||
version: "1.0"
|
||||
owner: Builder + Maintainer
|
||||
prerequisites: tests exist and pass, behavior is well-understood
|
||||
---
|
||||
|
||||
# Refactoring
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Code works but is hard to understand or maintain
|
||||
- Duplication exists that should be extracted
|
||||
- Naming is unclear or misleading
|
||||
- Structure doesn't match the conceptual model
|
||||
|
||||
## Core principle
|
||||
|
||||
**Refactoring changes structure, not behavior.** If you don't have tests, write
|
||||
them first. If you can't verify behavior is preserved, don't refactor.
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Verify baseline
|
||||
- Run the full test suite
|
||||
- Confirm all tests pass
|
||||
- Note the test output (baseline for comparison)
|
||||
|
||||
### 2. Identify the refactor
|
||||
- What specific structural improvement?
|
||||
- What is the expected benefit?
|
||||
- Is this the smallest useful refactor?
|
||||
|
||||
### 3. Make the change
|
||||
- One small step at a time
|
||||
- Run tests after each step
|
||||
- If tests fail, revert and try a smaller step
|
||||
|
||||
### 4. Verify
|
||||
- Run the full test suite again
|
||||
- Compare output to baseline
|
||||
- Verify no behavior change
|
||||
|
||||
### 5. Document
|
||||
- Note what was refactored and why
|
||||
- Update relevant documentation if public interfaces changed
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Refactoring while fixing a bug (mixes two changes)
|
||||
- Not having tests before refactoring
|
||||
- Making large changes in one step
|
||||
- Renaming things that don't need renaming
|
||||
- "While I'm here" scope creep
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Full test suite passes
|
||||
- No behavior change (same inputs, same outputs)
|
||||
- Code is clearer/simpler than before
|
||||
- Change is small enough to review
|
||||
@@ -0,0 +1,68 @@
|
||||
---
|
||||
name: repository-analysis
|
||||
description: Systematic repository exploration — understanding codebase structure, patterns, and conventions
|
||||
version: "1.0"
|
||||
owner: Explorer
|
||||
prerequisites: repo access
|
||||
---
|
||||
|
||||
# Repository Analysis
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- First encounter with a repository
|
||||
- Understanding a new area of a familiar repository
|
||||
- Before architectural or implementation decisions
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Orientation
|
||||
- Read README, package.json/pyproject.toml, or equivalent
|
||||
- Identify the project's purpose, language, and framework
|
||||
- Note the top-level structure
|
||||
|
||||
### 2. Entry points
|
||||
- Find the main entry points (main, index, app)
|
||||
- Trace the execution flow from entry to key functionality
|
||||
- Identify the public API surface
|
||||
|
||||
### 3. Structure
|
||||
- Map the directory structure to logical components
|
||||
- Identify module boundaries and dependencies
|
||||
- Note naming conventions
|
||||
|
||||
### 4. Build & test
|
||||
- Identify the build system and commands
|
||||
- Find the test suite and how to run it
|
||||
- Check for linting, type-checking, CI configuration
|
||||
|
||||
### 5. Conventions
|
||||
- Note coding style (formatting, naming, patterns)
|
||||
- Identify architectural patterns (MVC, layered, etc.)
|
||||
- Check for existing documentation of conventions
|
||||
|
||||
### 6. Dependencies
|
||||
- Review external dependencies
|
||||
- Note version constraints and lock files
|
||||
- Identify any custom or vendored dependencies
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Exploring too broadly (lost in the codebase)
|
||||
- Not recording findings (have to re-explore)
|
||||
- Confusing what exists with what is intended
|
||||
- Not distinguishing generated from hand-written code
|
||||
- Ignoring test files (they reveal intent)
|
||||
|
||||
## Evidence requirements
|
||||
|
||||
- File:line references for key findings
|
||||
- Confidence levels for uncertain findings
|
||||
- Questions for follow-up investigation
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- System map with key files and their roles
|
||||
- Confidence levels for each finding
|
||||
- Open questions listed
|
||||
- Findings recorded in report
|
||||
@@ -0,0 +1,57 @@
|
||||
---
|
||||
name: research
|
||||
description: Research methodology — structured investigation to answer questions
|
||||
version: "1.0"
|
||||
owner: Explorer + Writer
|
||||
prerequisites: question defined
|
||||
---
|
||||
|
||||
# Research
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- A question needs a well-researched answer
|
||||
- Comparing alternatives
|
||||
- Understanding best practices
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Define the question
|
||||
- What exactly are we trying to find out?
|
||||
- What constitutes a good answer?
|
||||
- What are the constraints?
|
||||
|
||||
### 2. Gather sources
|
||||
- Internal: code, docs, memory, agent reports
|
||||
- External: web, documentation sites, official references
|
||||
- Evaluate source reliability
|
||||
|
||||
### 3. Analyze
|
||||
- Compare findings across sources
|
||||
- Identify contradictions
|
||||
- Assess confidence in each finding
|
||||
|
||||
### 4. Synthesize
|
||||
- Combine findings into a coherent answer
|
||||
- Distinguish facts from opinions
|
||||
- State what remains uncertain
|
||||
|
||||
### 5. Record
|
||||
- Document findings with sources
|
||||
- Update memory if findings are durable
|
||||
- Share with relevant agents
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Accepting the first source without cross-referencing
|
||||
- Confusing opinions with facts
|
||||
- Not citing sources
|
||||
- Going too deep (research rabbit holes)
|
||||
- Not stopping when the question is answered
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Question answered with evidence
|
||||
- Sources cited
|
||||
- Confidence level stated
|
||||
- Uncertainties documented
|
||||
@@ -0,0 +1,70 @@
|
||||
---
|
||||
name: security-review
|
||||
description: Security review methodology — identifying and assessing security implications of changes
|
||||
version: "1.0"
|
||||
owner: Reviewer
|
||||
prerequisites: implementation completed, scope understood
|
||||
---
|
||||
|
||||
# Security Review
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Changes involving authentication, authorization, or access control
|
||||
- Changes to data handling, storage, or transmission
|
||||
- Changes to external API interactions
|
||||
- Changes to shell execution or system commands
|
||||
- Any change where security implications are uncertain
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
IDENTIFY CHANGE → ASSESS ATTACK SURFACE → CHECK VALIDATION → CHECK AUTH → CHECK DATA → CHECK DEPS → VERDICT
|
||||
```
|
||||
|
||||
## Security review checklist
|
||||
|
||||
### Input validation
|
||||
- [ ] All external inputs validated and sanitized
|
||||
- [ ] No SQL injection, command injection, or path traversal
|
||||
- [ ] File uploads validated (type, size, content)
|
||||
|
||||
### Authentication & authorization
|
||||
- [ ] Authentication checks are present and correct
|
||||
- [ ] Authorization checks enforce least privilege
|
||||
- [ ] No bypasses or backdoors
|
||||
|
||||
### Data protection
|
||||
- [ ] Sensitive data is not logged or exposed
|
||||
- [ ] Secrets are not hardcoded
|
||||
- [ ] Data at rest and in transit is protected appropriately
|
||||
|
||||
### Error handling
|
||||
- [ ] Errors do not leak sensitive information
|
||||
- [ ] Stack traces are not exposed to users
|
||||
- [ ] Graceful degradation on security failures
|
||||
|
||||
### Dependencies
|
||||
- [ ] Dependencies are from trusted sources
|
||||
- [ ] No known vulnerabilities in dependencies
|
||||
- [ ] Dependency versions are pinned
|
||||
|
||||
### Shell & system
|
||||
- [ ] Shell commands use safe execution patterns
|
||||
- [ ] File permissions are appropriate
|
||||
- [ ] Temporary files are handled securely
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Assuming "it's internal" means "it's safe"
|
||||
- Trusting user input without validation
|
||||
- Hardcoding credentials (even "temporary" ones)
|
||||
- Logging sensitive data
|
||||
- Not considering the attack surface
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- All checklist items addressed
|
||||
- Findings categorized: CRITICAL / HIGH / MEDIUM / LOW / INFO
|
||||
- Risk assessment for each finding
|
||||
- Remediation plan for non-INFO findings
|
||||
@@ -0,0 +1,79 @@
|
||||
---
|
||||
name: systematic-debugging
|
||||
description: Systematic debugging methodology — hypothesis-driven root cause investigation
|
||||
version: "1.0"
|
||||
owner: Detective
|
||||
prerequisites: symptom observed, reproducible failure available
|
||||
---
|
||||
|
||||
# Systematic Debugging
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- A test fails, a feature breaks, or behavior diverges from expectation
|
||||
- The root cause is not immediately obvious
|
||||
- Multiple possible causes exist
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
SYMPTOM → OBSERVE → HYPOTHESIZE → TEST → TRACE → ELIMINATE → ROOT CAUSE
|
||||
```
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Reproduce the symptom
|
||||
- Confirm you can trigger the failure reliably
|
||||
- Record the exact command, input, and output
|
||||
- Note the environment (OS, versions, state)
|
||||
|
||||
### 2. Observe the evidence
|
||||
- Read error messages carefully (every word matters)
|
||||
- Check logs, stack traces, and output
|
||||
- Note what changed since the last working state
|
||||
|
||||
### 3. Form hypotheses
|
||||
- List possible causes (aim for 3-5)
|
||||
- Rank by likelihood based on evidence
|
||||
- For each hypothesis: what would be true if this were the cause?
|
||||
|
||||
### 4. Test hypotheses
|
||||
- Design the cheapest test for the most likely hypothesis
|
||||
- Use binary elimination: each test should rule out at least one hypothesis
|
||||
- Record what you actually observed vs. what you expected
|
||||
|
||||
### 5. Trace the failure
|
||||
- Follow the execution path from symptom to cause
|
||||
- Add strategic print/log statements if needed
|
||||
- Narrow down: which component, which function, which line?
|
||||
|
||||
### 6. Eliminate alternatives
|
||||
- Explicitly state why other hypotheses are ruled out
|
||||
- Document the evidence for each elimination
|
||||
|
||||
### 7. Establish root cause
|
||||
- State the root cause with confidence level (high/medium/low)
|
||||
- Provide supporting evidence (file:line, command output)
|
||||
- Identify what conditions make this cause trigger
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Fixing the symptom without understanding the cause
|
||||
- Jumping to the most obvious hypothesis without testing it
|
||||
- Changing multiple things at once (can't isolate what fixed it)
|
||||
- Accepting "it works now" without understanding why
|
||||
- Not recording what was eliminated
|
||||
|
||||
## Evidence requirements
|
||||
|
||||
- Reproduction steps
|
||||
- Hypotheses with evidence for/against
|
||||
- Root cause statement with confidence level
|
||||
- Eliminated alternatives with reasoning
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Root cause identified with evidence
|
||||
- Alternative hypotheses eliminated with reasoning
|
||||
- Confidence level stated (high/medium/low)
|
||||
- Report handed off to appropriate agent for resolution
|
||||
@@ -0,0 +1,71 @@
|
||||
---
|
||||
name: tdd
|
||||
description: Test-Driven Development methodology — write tests first, implement to pass, refactor
|
||||
version: "1.0"
|
||||
owner: Builder + Tester
|
||||
prerequisites: test framework identified, build system working
|
||||
---
|
||||
|
||||
# Test-Driven Development
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Implementing new functionality where correctness can be verified
|
||||
- Bug fixes where a regression test should exist
|
||||
- Any change where the expected behavior is well-defined
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
RED → write a failing test that captures the requirement
|
||||
GREEN → implement the minimum code to make the test pass
|
||||
REFACTOR → improve the code while keeping all tests green
|
||||
```
|
||||
|
||||
## Step-by-step procedure
|
||||
|
||||
### 1. Understand the requirement
|
||||
- What behavior is expected?
|
||||
- What inputs produce what outputs?
|
||||
- What edge cases exist?
|
||||
- What error conditions should be handled?
|
||||
|
||||
### 2. Write the test (RED)
|
||||
- Write the smallest test that captures one aspect of the requirement
|
||||
- Verify the test fails for the right reason (not a syntax error)
|
||||
- Run the test to confirm it fails
|
||||
|
||||
### 3. Implement (GREEN)
|
||||
- Write the minimum code to make the test pass
|
||||
- Do not add behavior not captured by a test
|
||||
- Run the test to confirm it passes
|
||||
|
||||
### 4. Verify (REFACTOR)
|
||||
- Run the full test suite (not just the new test)
|
||||
- Refactor if needed: extract methods, rename for clarity, remove duplication
|
||||
- Verify tests still pass after each refactor step
|
||||
|
||||
### 5. Repeat
|
||||
- Return to step 1 for the next aspect of the requirement
|
||||
- Stop when all aspects are covered and tests pass
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Writing tests after implementation (loses the design benefit)
|
||||
- Writing too many tests at once (harder to isolate failures)
|
||||
- Testing implementation details instead of behavior
|
||||
- Skipping the refactor step (technical debt accumulates)
|
||||
- Not running the full suite after changes
|
||||
|
||||
## Evidence requirements
|
||||
|
||||
- Test file with the new test(s)
|
||||
- Test output showing RED → GREEN progression
|
||||
- Full suite passing after completion
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- All new tests pass
|
||||
- Full existing suite passes
|
||||
- Tests capture the behavior, not the implementation
|
||||
- Each test is independently runnable
|
||||
@@ -0,0 +1,59 @@
|
||||
---
|
||||
name: test-analysis
|
||||
description: Test analysis — evaluating test quality, coverage gaps, and test strategy
|
||||
version: "1.0"
|
||||
owner: Tester
|
||||
prerequisites: test suite exists
|
||||
---
|
||||
|
||||
# Test Analysis
|
||||
|
||||
## When to use this skill
|
||||
|
||||
- Evaluating test quality before accepting changes
|
||||
- Identifying coverage gaps
|
||||
- Designing test strategy for new features
|
||||
|
||||
## Core methodology
|
||||
|
||||
```text
|
||||
EXAMINE TESTS → ASSESS COVERAGE → EVALUATE QUALITY → IDENTIFY GAPS → PRIORITIZE → RECOMMEND
|
||||
```
|
||||
|
||||
## Analysis dimensions
|
||||
|
||||
### Coverage
|
||||
- What code paths are exercised?
|
||||
- What branches are tested?
|
||||
- What error conditions are covered?
|
||||
|
||||
### Quality
|
||||
- Do tests verify behavior (not implementation)?
|
||||
- Are tests independent (no ordering dependencies)?
|
||||
- Are tests deterministic (no flakiness)?
|
||||
- Are assertions meaningful (not just "no crash")?
|
||||
|
||||
### Completeness
|
||||
- Are edge cases tested?
|
||||
- Are boundary conditions covered?
|
||||
- Are error paths tested?
|
||||
- Are integration points verified?
|
||||
|
||||
### Maintainability
|
||||
- Are tests readable?
|
||||
- Are tests well-organized?
|
||||
- Are tests fast enough for the feedback loop?
|
||||
|
||||
## Common pitfalls
|
||||
|
||||
- Testing implementation details (breaks on refactor)
|
||||
- Writing tests that always pass (no real verification)
|
||||
- Missing the failure path (happy path only)
|
||||
- Slow tests that discourage running them
|
||||
- Flaky tests that erode confidence
|
||||
|
||||
## Exit criteria
|
||||
|
||||
- Coverage gaps identified with severity
|
||||
- Test quality assessment complete
|
||||
- Recommendations for improvement prioritized
|
||||
Reference in New Issue
Block a user