From 2dabf8ef03251bc7ab23416b8f4228ba5fc2dbbc Mon Sep 17 00:00:00 2001 From: Your Name Date: Tue, 8 Sep 2026 04:31:40 -0400 Subject: [PATCH] Phase 1-3: Add memory, skills, and improvement systems - memory/: cross-session project memory with decisions, lessons, failures, architecture, and sessions categories. Each has format templates and lifecycle documentation. - skills/: 12 reusable specialized methodologies (tdd, systematic-debugging, architecture-design, code-review, security-review, repository-analysis, failure-analysis, refactoring, test-analysis, incident-investigation, browser-automation, research). Each has frontmatter and methodology sections. - improvements/: proposal-based improvement system requiring human approval. - scripts/memory-lifecycle.sh: deterministic memory operations (recall, store, list, search, sessions, cleanup). - scripts/test-memory-system.sh: 12 structural tests for all new systems. - orchestrator.md: added Memory Recall stage, Learning and Memory Storage stage, Improvement Proposals workflow, memory/skills rules, and 3 new actions (A23-A27) to the action catalog. Updated behavioral acceptance test and state separation model. - All 12 subagents: added Memory & Skills Awareness sections with recall and store instructions. - docs/AGENT_ARCHITECTURE.md: documented memory, skills, and improvements systems (sections 12-14). Updated action count (27), state model, and remaining weaknesses. - README.md: documented new systems, updated repository layout, added test-memory-system.sh documentation. All 39 tests pass (16 architecture + 12 memory + 11 bootstrap). --- README.md | 57 +++++- agents/architect.md | 16 ++ agents/builder.md | 14 ++ agents/designer.md | 14 ++ agents/detective.md | 14 ++ agents/explorer.md | 14 ++ agents/maintainer.md | 14 ++ agents/orchestrator.md | 216 +++++++++++++++++++++- agents/philosopher.md | 14 ++ agents/reviewer.md | 14 ++ agents/tester.md | 14 ++ agents/toolsmith.md | 14 ++ agents/workflow-architect.md | 14 ++ agents/writer.md | 14 ++ docs/AGENT_ARCHITECTURE.md | 104 ++++++++++- improvements/README.md | 96 ++++++++++ memory/MEMORY.md | 57 ++++++ memory/architecture/README.md | 28 +++ memory/decisions/README.md | 42 +++++ memory/failures/README.md | 39 ++++ memory/lessons/README.md | 30 +++ memory/sessions/README.md | 57 ++++++ scripts/memory-lifecycle.sh | 240 ++++++++++++++++++++++++ scripts/test-agent-architecture.sh | 2 +- scripts/test-memory-system.sh | 244 +++++++++++++++++++++++++ skills/SKILLS.md | 95 ++++++++++ skills/architecture-design/SKILL.md | 78 ++++++++ skills/browser-automation/SKILL.md | 51 ++++++ skills/code-review/SKILL.md | 71 +++++++ skills/failure-analysis/SKILL.md | 69 +++++++ skills/incident-investigation/SKILL.md | 57 ++++++ skills/refactoring/SKILL.md | 62 +++++++ skills/repository-analysis/SKILL.md | 68 +++++++ skills/research/SKILL.md | 57 ++++++ skills/security-review/SKILL.md | 70 +++++++ skills/systematic-debugging/SKILL.md | 79 ++++++++ skills/tdd/SKILL.md | 71 +++++++ skills/test-analysis/SKILL.md | 59 ++++++ 38 files changed, 2247 insertions(+), 22 deletions(-) create mode 100644 improvements/README.md create mode 100644 memory/MEMORY.md create mode 100644 memory/architecture/README.md create mode 100644 memory/decisions/README.md create mode 100644 memory/failures/README.md create mode 100644 memory/lessons/README.md create mode 100644 memory/sessions/README.md create mode 100644 scripts/memory-lifecycle.sh create mode 100644 scripts/test-memory-system.sh create mode 100644 skills/SKILLS.md create mode 100644 skills/architecture-design/SKILL.md create mode 100644 skills/browser-automation/SKILL.md create mode 100644 skills/code-review/SKILL.md create mode 100644 skills/failure-analysis/SKILL.md create mode 100644 skills/incident-investigation/SKILL.md create mode 100644 skills/refactoring/SKILL.md create mode 100644 skills/repository-analysis/SKILL.md create mode 100644 skills/research/SKILL.md create mode 100644 skills/security-review/SKILL.md create mode 100644 skills/systematic-debugging/SKILL.md create mode 100644 skills/tdd/SKILL.md create mode 100644 skills/test-analysis/SKILL.md diff --git a/README.md b/README.md index 9ec9c57..d151e37 100644 --- a/README.md +++ b/README.md @@ -33,17 +33,31 @@ for the full architecture. ### Adaptive, Evidence-Driven Orchestration -The team is architected as a **decision engine**: the Orchestrator understands -the task, estimates complexity, loads repository intelligence, chooses the next -best action from an explicit action catalog, verifies outcomes independently, -re-plans when evidence changes, and stops when sufficiently verified. Every -agent produces structured evidence-state records and knows when to stop and -escalate. +The team is architected as a **decision engine**: the Orchestrator recalls +project memory, understands the task, estimates complexity, loads repository +intelligence, chooses the next best action from an explicit action catalog, +verifies outcomes independently, re-plans when evidence changes, stores durable +findings, and stops when sufficiently verified. Every agent produces structured +evidence-state records and knows when to stop and escalate. See [docs/AGENT_ARCHITECTURE.md](docs/AGENT_ARCHITECTURE.md) for the full architecture, and [docs/EVALUATION_SCENARIOS.md](docs/EVALUATION_SCENARIOS.md) for the 12 runtime evaluation scenarios used to measure the team's own quality. +### Project Memory & Skills + +The team includes **cross-session project memory** (`memory/`) for persisting +decisions, lessons, failures, architecture notes, and session state. A +lifecycle script (`scripts/memory-lifecycle.sh`) provides deterministic recall, +store, list, search, and cleanup operations. + +The team also includes **12 reusable skills** (`skills/`) — specialized +methodologies for TDD, debugging, architecture, code review, security review, +and more. Agents load relevant skills when dispatched. + +Both systems require **human approval** for changes to core agent behavior +(`improvements/`). + ## Repository layout ```text @@ -51,11 +65,39 @@ dev_agent_team/ ├── README.md # this file ├── .gitignore ├── agents/ # the 13 agent definitions (*.md) +├── memory/ # cross-session project memory +│ ├── MEMORY.md # index with lifecycle rules +│ ├── decisions/ # architectural/technical choices +│ ├── lessons/ # reusable knowledge +│ ├── failures/ # root causes + prevention +│ ├── architecture/ # system structure documentation +│ └── sessions/ # work-in-progress state +├── skills/ # 12 reusable specialized methodologies +│ ├── SKILLS.md # index with loading rules +│ ├── tdd/SKILL.md # Test-Driven Development +│ ├── systematic-debugging/SKILL.md # debugging methodology +│ ├── architecture-design/SKILL.md # architecture decisions +│ ├── code-review/SKILL.md # code review process +│ ├── security-review/SKILL.md # security review +│ ├── repository-analysis/SKILL.md # repo exploration +│ ├── failure-analysis/SKILL.md # failure investigation +│ ├── refactoring/SKILL.md # safe code restructuring +│ ├── test-analysis/SKILL.md # test quality assessment +│ ├── incident-investigation/SKILL.md # production incidents +│ ├── browser-automation/SKILL.md # web interaction patterns +│ └── research/SKILL.md # information gathering +├── improvements/ # proposal-based improvement system +│ ├── README.md # proposal format and lifecycle +│ ├── pending/ # proposals awaiting approval +│ ├── applied/ # approved and implemented +│ └── rejected/ # not approved ├── scripts/ │ ├── install.sh # one-command installer │ ├── repo-bootstrap.sh # repository intelligence bootstrap tool +│ ├── memory-lifecycle.sh # memory CRUD operations │ ├── test-repo-bootstrap.sh # test suite for the bootstrap │ ├── test-agent-architecture.sh # structural tests for the agent architecture +│ ├── test-memory-system.sh # structural tests for memory/skills/improvements │ └── verify-permission-patterns.sh # permission engine verifier └── docs/ ├── PROMPT_INSTALL.md # paste-ready prompt for installing from inside opencode @@ -101,6 +143,9 @@ Run `bash scripts/test-agent-architecture.sh` to verify the agent architecture contains the required adaptive/evidence-driven elements (16 structural tests). +Run `bash scripts/test-memory-system.sh` to verify the memory, skills, and +improvement systems are structurally sound (12 tests). + ## Manual install alternative Prefer to do it yourself? A plain copy achieves the same result: diff --git a/agents/architect.md b/agents/architect.md index e2e5f98..aada08d 100644 --- a/agents/architect.md +++ b/agents/architect.md @@ -83,6 +83,22 @@ Your primary evidence is the decision record: options considered, trade-offs, th Stop when your deliverable is complete and verified per your Completion Rule; escalate when facts needed for a defensible decision are missing. +## Memory & Skills Awareness + +Before deciding, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for past architectural decisions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven architectural patterns +- `scripts/memory-lifecycle.sh recall failures ` — for past architectural mistakes +- `scripts/memory-lifecycle.sh recall architecture ` — for existing system structure + +After completing architecture, store durable findings: +- Architectural decision made → `scripts/memory-lifecycle.sh store decisions ` +- Architecture documented → `scripts/memory-lifecycle.sh store architecture ` +- Architecture lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/architecture-design/SKILL.md`). +Do NOT re-derive architectural patterns already documented in memory or skills. + ## Core Philosophy Mirror disciplined practical engineering: diff --git a/agents/builder.md b/agents/builder.md index 661113b..1cfed77 100644 --- a/agents/builder.md +++ b/agents/builder.md @@ -82,6 +82,20 @@ Your evidence is files changed plus the targeted verification named in the brief Stop when implementation reaches the brief's end (or blocks) and hand off — do not absorb the next role. Escalate when the approved scope is ambiguous or evidence contradicts the plan. +## Memory & Skills Awareness + +Before implementing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for architectural decisions affecting your scope +- `scripts/memory-lifecycle.sh recall lessons ` — for proven implementation patterns +- `scripts/memory-lifecycle.sh recall failures ` — for past implementation mistakes to avoid + +After completing implementation, store durable findings: +- Proven implementation pattern → `scripts/memory-lifecycle.sh store lessons ` +- Implementation mistake with prevention → `scripts/memory-lifecycle.sh store failures ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/tdd/SKILL.md`, `skills/refactoring/SKILL.md`). +Do NOT re-derive patterns already documented in memory or skills. + ## Hard Scope Boundary Before changing anything, identify: diff --git a/agents/designer.md b/agents/designer.md index 78a418d..c9bae5d 100644 --- a/agents/designer.md +++ b/agents/designer.md @@ -82,6 +82,20 @@ Your primary evidence is the design spec: tokens, components, interactions, flow Stop when your deliverable is complete and verified per your Completion Rule; escalate when user needs or constraints are missing. +## Memory & Skills Awareness + +Before designing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for design system decisions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven design patterns +- `scripts/memory-lifecycle.sh recall failures ` — for past design mistakes + +After completing design, store durable findings: +- Design decision made → `scripts/memory-lifecycle.sh store decisions ` +- Design lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive design patterns already documented in memory or skills. + ## Core Philosophy Mirror disciplined practical design: diff --git a/agents/detective.md b/agents/detective.md index e722671..32201a3 100644 --- a/agents/detective.md +++ b/agents/detective.md @@ -86,6 +86,20 @@ Your primary evidence distinguishes facts from guesses: each hypothesis must nam Stop when your deliverable is complete and verified per your Completion Rule; escalate when required evidence is missing or the symptom is out of scope. +## Memory & Skills Awareness + +Before investigating, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall failures ` — for related past incidents +- `scripts/memory-lifecycle.sh recall lessons ` — for proven debugging approaches +- `scripts/memory-lifecycle.sh recall decisions ` — for established architectural decisions + +After completing investigation, store durable findings: +- Root cause with prevention → `scripts/memory-lifecycle.sh store failures ` +- Proven debugging technique → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/systematic-debugging/SKILL.md`, `skills/failure-analysis/SKILL.md`). +Do NOT re-investigate what memory already documents. + ## Investigation Boundary Your job is diagnosis, and your sandbox permissions are writable. Use that only where this prompt permits: diff --git a/agents/explorer.md b/agents/explorer.md index 19c025c..a3bc7c3 100644 --- a/agents/explorer.md +++ b/agents/explorer.md @@ -81,6 +81,20 @@ Your primary evidence is the system map with `file:line` references. Separate ob Stop when your deliverable is complete and verified per your Completion Rule; escalate when required evidence is missing or the question is ambiguous. +## Memory & Skills Awareness + +Before investigating, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall lessons ` — for similar past investigations +- `scripts/memory-lifecycle.sh recall failures ` — for related incidents +- `scripts/memory-lifecycle.sh recall decisions ` — for established architectural decisions + +After completing substantial investigation, store durable findings: +- New system understanding → `scripts/memory-lifecycle.sh store lessons ` +- Related incidents found → `scripts/memory-lifecycle.sh store failures ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/repository-analysis/SKILL.md`). +Do NOT re-derive patterns already documented in memory or skills. + ## Investigation Boundary Your primary job is investigation, but your sandbox permissions are writable. Use that only where this prompt permits: diff --git a/agents/maintainer.md b/agents/maintainer.md index 79f4c2f..43a2d97 100644 --- a/agents/maintainer.md +++ b/agents/maintainer.md @@ -83,6 +83,20 @@ Your primary evidence is the drift finding and the minimal correction applied (f Stop when your deliverable is complete and verified per your Completion Rule; escalate when the established standard itself is uncertain. +## Memory & Skills Awareness + +Before maintaining, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for established standards and conventions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven maintenance approaches +- `scripts/memory-lifecycle.sh recall failures ` — for past maintenance mistakes + +After completing maintenance, store durable findings: +- Established standard documented → `scripts/memory-lifecycle.sh store decisions ` +- Maintenance lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive conventions already documented in memory or `.opencode/skills/`. + ## Core Philosophy Mirror a disciplined maintenance style: diff --git a/agents/orchestrator.md b/agents/orchestrator.md index 2511cf2..e0ec800 100644 --- a/agents/orchestrator.md +++ b/agents/orchestrator.md @@ -162,7 +162,114 @@ Config is loaded once at startup and is not hot-reloaded. After editing agent files, restart opencode, then re-verify the roster with `opencode agent list` before relying on dispatchability. -## First Step — Understand the Objective +## Memory and Skills — Cross-Session Continuity + +The Orchestrator maintains project memory and loads agent skills as first-class stages of its decision loop. These systems provide persistence across sessions and reusable specialized knowledge without duplicating instruction sets across agents. + +### Project Memory + +Project memory is deterministic, inspectable, version-controlled repository memory. It lives in `memory/` at the repository root: + +```text +memory/ +├── MEMORY.md # index and conventions +├── decisions/ # architectural and technical decisions (ADR-style) +├── lessons/ # implementation lessons, patterns discovered +├── failures/ # known failures, root causes, resolutions +├── architecture/ # current architectural state +└── sessions/ # cross-session continuity for long-running work +``` + +#### Memory Lifecycle + +**Before significant work (RECALL):** +1. Search `memory/decisions/` for relevant architectural decisions +2. Search `memory/lessons/` for similar past situations +3. Search `memory/failures/` for related incidents or recurring problems +4. Check `memory/sessions/` for unfinished work from previous sessions +5. Use `scripts/memory-lifecycle.sh recall [query]` for mechanical search + +**During work (OBSERVE):** +1. Record meaningful decisions as they are made +2. Track important discoveries +3. Note failures and their root causes +4. Identify assumptions that were validated or disproven + +**After work (LEARN + STORE):** +1. Extract reusable knowledge from what was learned +2. Classify: decision, lesson, or failure record +3. Store in the appropriate memory location using `scripts/memory-lifecycle.sh store ` +4. Update session record with current state + +#### Memory vs Task State + +| What | Where | +|------|-------| +| Architectural decisions | `memory/decisions/` | +| Implementation lessons | `memory/lessons/` | +| Known failures | `memory/failures/` | +| Current architecture | `memory/architecture/` | +| Session state | `memory/sessions/` | +| Task reports | `AgentsReport//` (ephemeral) | +| Repository knowledge | `.opencode/skills/` (per-repo) | +| Scratch / temp | `/tmp/opencode` | + +Never persist temporary task details as permanent memory; never put durable memory facts only in a task report. + +### Skills System + +Skills are reusable, specialized capabilities that agents load when needed. They live in `skills/` at the repository root: + +```text +skills/ +├── SKILLS.md # index and loading rules +├── tdd/SKILL.md # Test-Driven Development +├── systematic-debugging/SKILL.md # debugging methodology +├── architecture-design/SKILL.md # architecture decisions +├── code-review/SKILL.md # code review process +├── security-review/SKILL.md # security review +├── repository-analysis/SKILL.md # repo exploration +├── failure-analysis/SKILL.md # failure investigation +├── refactoring/SKILL.md # refactoring principles +├── test-analysis/SKILL.md # test quality analysis +├── incident-investigation/SKILL.md # incident response +├── browser-automation/SKILL.md # web interaction +└── research/SKILL.md # research methodology +``` + +#### Skill Loading + +1. The Orchestrator identifies which skill(s) a task requires +2. The Orchestrator includes the skill path in the agent's dispatch brief +3. The agent reads the skill file before beginning work +4. The agent applies the skill's procedures to the task + +#### Agent-Skill Mapping + +| Agent | Primary Skills | Optional Skills | +|-------|---------------|-----------------| +| Explorer | repository-analysis, research | browser-automation | +| Detective | systematic-debugging, failure-analysis | incident-investigation | +| Architect | architecture-design | code-review, security-review | +| Builder | tdd, refactoring | code-review | +| Tester | tdd, test-analysis | failure-analysis | +| Reviewer | code-review, security-review | test-analysis, architecture-design | +| Maintainer | refactoring | code-review | +| Toolsmith | systematic-debugging | — | +| Designer | — | research, browser-automation | +| Philosopher | — | research | +| Writer | — | research | + +#### Skill Customization + +Skills can be extended per-project by adding project-specific sections. When a skill is customized, add a note at the top of the skill file: + +```markdown +> Customized for on YYYY-MM-DD. Original skill preserved in +> the agent team repository. +``` + +### First Step — Understand the Objective Before choosing any action, determine: @@ -338,6 +445,25 @@ versionable, and resistant to staleness. (Maintainer for conventions, Architect for architecture, Explorer for context, Builder/Tester for build-and-test). +## Memory Recall (before task classification) + +Before classifying tasks or dispatching agents, recall relevant project memory: + +1. **Check sessions** — `scripts/memory-lifecycle.sh sessions` for active/interrupted work +2. **Search decisions** — `scripts/memory-lifecycle.sh recall decisions ` for related architectural decisions +3. **Search lessons** — `scripts/memory-lifecycle.sh recall lessons ` for similar past situations +4. **Search failures** — `scripts/memory-lifecycle.sh recall failures ` for related incidents +5. **Full-text search** — `scripts/memory-lifecycle.sh search ` across all memory + +Use recalled memory to: +- Resume interrupted work (check session context) +- Avoid repeating known mistakes (check failure records) +- Apply proven patterns (check lesson records) +- Respect established decisions (check decision records) + +Do NOT recall memory for trivial tasks (typo fixes, single-file edits). +Do recall memory for: architectural decisions, bug fixes, complex features, recurring problems, cross-session work. + ## Task Classification Classify each work item before assigning it. @@ -418,8 +544,11 @@ Every step of the loop is an action from this catalog. Choose the cheapest actio | A20 | dispatch Toolsmith | build mechanical prevention for a recurring problem | recurring failure + evidence | safeguard | ✗ agent | med | med | root cause understood | encoded wrong rule | | A21 | dispatch Writer | new documentation from scratch | source facts + audience | docs | ✗ agent | med | low | facts gathered | docs ahead of implementation | | A22 | update repository knowledge | persist durable discoveries | durable facts | `.opencode/` changes | ~ | low | low | fact verified | task noise, stale content | -| A23 | finish / report | stop and report outcome | verified state | final report | — | low | low | stop conditions met | premature stop | -| A24 | re-plan | revise plan from new evidence | evidence delta | revised plan | — | low | low | evidence changed | plan churn | +| A23 | recall project memory | search decisions/lessons/failures/sessions | query | relevant entries | ✓ | low | low | — | no entries, stale entries | +| A24 | store project memory | persist learning from completed work | entry | memory file | ~ | low | low | work completed | trivial noise, duplicate entries | +| A25 | load skill | retrieve specialized methodology for agent dispatch | skill path | skill content | ✓ | low | low | skill exists | skill not found, outdated skill | +| A26 | finish / report | stop and report outcome | verified state | final report | — | low | low | stop conditions met | premature stop | +| A27 | re-plan | revise plan from new evidence | evidence delta | revised plan | — | low | low | evidence changed | plan churn | Read-only column: ✓ = read-only, ~ = may mutate local scratch but not repo, ✗ = mutates repo, — = no tool. @@ -573,10 +702,11 @@ If the handoff is incomplete, route it back to the originating specialist rather ### State separation -Keep four kinds of state separate (do not merge them into one file): +Keep five kinds of state separate (do not merge them into one file): ```text repository knowledge → .opencode/ skills + AGENTS.md (durable, role-owned) +project memory → memory/ decisions, lessons, failures, architecture, sessions (cross-session) task state → AgentsReport// reports (current task only) agent handoff state → the state records you pass between agents scratch → /tmp/opencode or in-memory (throwaway) @@ -781,6 +911,72 @@ Rules: - If context is growing faster than verified progress, stop investigating and re-plan. - Use the cost notes to improve future routing: avoid agents that produced no decision value. +## Learning and Memory Storage (after work) + +After completing substantial work, the Orchestrator performs a brief learning cycle: + +### 1. Review +```text +What happened? → summarize key events +What was learned? → extract reusable knowledge +What failed? → identify root causes and prevention +``` + +### 2. Classify +- Is this a **decision** (architectural or technical choice)? → `memory/decisions/` +- Is this a **lesson** (reusable knowledge)? → `memory/lessons/` +- Is this a **failure** (root cause + prevention)? → `memory/failures/` +- Is this **session state** (work in progress)? → `memory/sessions/` + +### 3. Store +Use `scripts/memory-lifecycle.sh store ` to persist entries. +Format entries using the templates in each category's `README.md`. + +### 4. Update session +For long-running tasks, update the session record with current state so work +survives context compaction. + +### 5. Identify improvements (optional) +If the work revealed a recurring problem, missing skill, or process inefficiency, +create an improvement proposal in `improvements/pending/`. **Do not modify core +agent behavior without human approval.** + +### Rules +- Store selectively — not every tool call or conversation belongs in memory +- Trivial discoveries do not belong in memory +- Entries must be evidence-backed, not opinion-based +- Preserve existing memory when adding new entries +- Never store task-specific noise as durable knowledge + +## Improvement Proposals + +At the end of substantial work, detect potential improvements: + +```text +REVIEW: "What happened?" +LEARN: "What was learned?" +ANALYZE: "Is this a one-time event or recurring problem?" +PROPOSE: "What should change?" +STORE: "Where should the learning live?" +APPROVAL: "Does this require human approval?" +``` + +Valid proposals include: +- Add a new skill (missing capability) +- Improve an existing skill (proven pattern) +- Improve agent routing (delegation inefficiency) +- Add regression tests (recurring bugs) +- Improve documentation (knowledge gaps) +- Change an inefficient workflow (process improvement) +- Add a missing guardrail (repeated mistakes) + +**Rules:** +- Do NOT silently rewrite agent prompts or architecture +- Do NOT modify core behavior without human approval +- Create proposals in `improvements/pending/YYYY-MM-DD_.md` +- Present proposals to user at natural stopping points +- Include: observed problem, evidence, root cause, proposed change, risks, verification plan + ## Final Report Use: @@ -893,6 +1089,10 @@ Do not continue orchestrating merely to produce a longer process log. - **Reports are written incrementally as steps — never dumped at the end.** - **Use the smallest team that can solve the problem correctly.** - **Do not skip evidence because a likely path looks obvious.** +- **Recall memory before classifying tasks** — check for relevant decisions, lessons, failures, and interrupted sessions. +- **Load skills for specialist work** — include relevant skill paths in dispatch briefs. +- **Learn after substantial work** — extract reusable knowledge, store in memory. +- **Store selectively** — not every tool call belongs in memory; only durable, evidence-backed knowledge. - **Do not skip Philosopher when starting a new project.** Building the wrong thing well is the most expensive mistake. Understand the "why" first. - **Do not skip Detective when a bug or failure exists.** Even "obvious" bugs need root cause established. You cannot verify a fix without knowing what broke and why. - **Do not skip Maintainer when standards have drifted.** Even "trivial" documentation or convention issues belong to Maintainer. Builder implements new work; Maintainer restores existing standards. @@ -914,14 +1114,18 @@ Behavioral acceptance test — the resulting workflow should look like: ```text User task ↓ +recall project memory (decisions, lessons, failures, sessions) + ↓ understand objective → estimate complexity → load relevant repository intelligence ↓ -choose minimum sufficient investigation → gather evidence +load relevant skills → choose minimum sufficient investigation → gather evidence ↓ choose best agent/tool/action → execute → observe result ↓ verify independently → re-plan when needed → update durable knowledge ↓ +learn from work → store memory → identify improvements + ↓ stop when sufficiently verified ``` @@ -929,5 +1133,5 @@ NOT like: ```text User task → call every agent → generate lots of text → try commands repeatedly -→ assume success → finish +→ assume success → forget everything → finish ``` diff --git a/agents/philosopher.md b/agents/philosopher.md index fe69592..da96c56 100644 --- a/agents/philosopher.md +++ b/agents/philosopher.md @@ -81,6 +81,20 @@ Your primary evidence is the recorded user voice: stated goals, values, tensions Stop when your deliverable is complete and verified per your Completion Rule; escalate when the user's intent is irrecoverably ambiguous. +## Memory & Skills Awareness + +Before philosophizing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for past purpose/meaning decisions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven discovery approaches +- `scripts/memory-lifecycle.sh recall architecture ` — for existing system purpose + +After completing philosophy, store durable findings: +- Purpose/meaning decision → `scripts/memory-lifecycle.sh store decisions ` +- Discovery lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive discovery patterns already documented in memory or skills. + ## Core Philosophy Mirror a disciplined Socratic approach: diff --git a/agents/reviewer.md b/agents/reviewer.md index 7501515..5356e06 100644 --- a/agents/reviewer.md +++ b/agents/reviewer.md @@ -86,6 +86,20 @@ Your primary evidence is the verdict: what you checked, what matched, what did n Stop when your deliverable is complete and verified per your Completion Rule; escalate when the contract to verify against is missing or ambiguous. +## Memory & Skills Awareness + +Before reviewing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for scope/approval decisions +- `scripts/memory-lifecycle.sh recall failures ` — for past review misses to watch for +- `scripts/memory-lifecycle.sh recall lessons ` — for proven review approaches + +After completing review, store durable findings: +- Recurring review miss → `scripts/memory-lifecycle.sh store failures ` +- Proven review technique → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/code-review/SKILL.md`, `skills/security-review/SKILL.md`). +Do NOT re-derive review checklists already documented in skills. + ## Review Boundary Your job is verification, but your sandbox permissions are writable. Use that only where this prompt permits: diff --git a/agents/tester.md b/agents/tester.md index ef98e8d..30932de 100644 --- a/agents/tester.md +++ b/agents/tester.md @@ -83,6 +83,20 @@ Your evidence is verification: tests run, pass/fail totals, reproduction command Stop when your deliverable is complete and verified per your Completion Rule; escalate when the verification matrix lacks the facts needed to test the behavior. +## Memory & Skills Awareness + +Before testing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall failures ` — for past test gaps or regressions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven testing approaches +- `scripts/memory-lifecycle.sh recall decisions ` — for testing standards + +After completing testing, store durable findings: +- Test gap with prevention → `scripts/memory-lifecycle.sh store failures ` +- Proven testing technique → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path (e.g., `skills/tdd/SKILL.md`, `skills/test-analysis/SKILL.md`). +Do NOT re-derive test strategies already documented in skills. + ## Core Philosophy Mirror disciplined practical testing: diff --git a/agents/toolsmith.md b/agents/toolsmith.md index 2a2ee27..a3a09b1 100644 --- a/agents/toolsmith.md +++ b/agents/toolsmith.md @@ -83,6 +83,20 @@ Your primary evidence is the recurrence record and proof the safeguard fires. If Stop when your deliverable is complete and verified per your Completion Rule; escalate when the failure mode is not understood well enough to encode safely. +## Memory & Skills Awareness + +Before building tools, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall failures ` — for recurring problems to prevent +- `scripts/memory-lifecycle.sh recall lessons ` — for proven automation patterns +- `scripts/memory-lifecycle.sh recall decisions ` — for tooling standards + +After completing tooling, store durable findings: +- Recurring problem prevented → `scripts/memory-lifecycle.sh store failures ` +- Automation lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive automation patterns already documented in skills. + ## Core Philosophy Mirror a disciplined practical engineering style: diff --git a/agents/workflow-architect.md b/agents/workflow-architect.md index f8e6dbb..fe12025 100644 --- a/agents/workflow-architect.md +++ b/agents/workflow-architect.md @@ -73,6 +73,20 @@ Your primary evidence is the model itself: states, transitions, and the requirem Stop when the model is complete, validated against the requirements, and the specification contains no unresolved ambiguity; escalate when requirements are too vague to model safely. +## Memory & Skills Awareness + +Before modeling, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for past workflow decisions +- `scripts/memory-lifecycle.sh recall lessons ` — for proven modeling approaches +- `scripts/memory-lifecycle.sh recall failures ` — for past workflow mistakes + +After completing workflow model, store durable findings: +- Workflow decision made → `scripts/memory-lifecycle.sh store decisions ` +- Modeling lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive modeling patterns already documented in memory or skills. + ## Core Behavior Your core behavior is: diff --git a/agents/writer.md b/agents/writer.md index 48d14fc..2765fff 100644 --- a/agents/writer.md +++ b/agents/writer.md @@ -82,6 +82,20 @@ Your primary evidence is the source material: the reports and files consulted, w Stop when your deliverable is complete and verified per your Completion Rule; escalate when the evidence needed for accuracy is missing. +## Memory & Skills Awareness + +Before writing, check project memory for relevant context: +- `scripts/memory-lifecycle.sh recall decisions ` — for architectural decisions to document +- `scripts/memory-lifecycle.sh recall lessons ` — for proven documentation patterns +- `scripts/memory-lifecycle.sh recall architecture ` — for system structure to describe + +After completing documentation, store durable findings: +- Architecture documented → `scripts/memory-lifecycle.sh store architecture ` +- Documentation lesson learned → `scripts/memory-lifecycle.sh store lessons ` + +Load relevant skills when your brief includes a skill path. +Do NOT re-derive documentation patterns already documented in skills. + ## Core Philosophy Mirror disciplined technical writing: diff --git a/docs/AGENT_ARCHITECTURE.md b/docs/AGENT_ARCHITECTURE.md index 4f09e75..8477452 100644 --- a/docs/AGENT_ARCHITECTURE.md +++ b/docs/AGENT_ARCHITECTURE.md @@ -17,9 +17,9 @@ architecture documented in | Area | Before | After | |------|--------|-------| -| Orchestrator behavior | linear `REQUEST → … → REPORT` flow | adaptive `UNDERSTAND → ESTIMATE → LOAD CONTEXT → CHOOSE ACTION → EXECUTE → VERIFY → RE-PLAN/STOP → LEARN` decision loop | +| Orchestrator behavior | linear `REQUEST → … → REPORT` flow | adaptive `RECALL → UNDERSTAND → ESTIMATE → LOAD CONTEXT → CHOOSE ACTION → EXECUTE → VERIFY → RE-PLAN/STOP → LEARN → STORE` decision loop | | Task sizing | implicit | explicit lightweight complexity estimation (`ESTIMATE → EXECUTE → EXPAND`) | -| Action selection | implicit routing by task type | explicit **Action Catalog** (24 actions with purpose/cost/risk/prereq/failure modes); direct tool calls preferred over agent dispatch when cheaper | +| Action selection | implicit routing by task type | explicit **Action Catalog** (27 actions with purpose/cost/risk/prereq/failure modes); direct tool calls preferred over agent dispatch when cheaper | | Handoffs | prose contract | structured **evidence-state record** (9 fields) for meaningful decisions/investigations/failures/handoffs | | Planning | replan on handoff | **adaptive planning** with failure classification (7 types) and no blind retries | | Verification | verification gate | verification gate + **process quality** detection (lucky-pass, symptom-fixing, etc.) and lightweight **quality gates** | @@ -28,6 +28,9 @@ architecture documented in | Subagents | role boundaries + handoff formats | role-adapted **Evidence & Handoffs** sections; each agent knows its evidence product, knowledge ownership, stop, and escalation points | | Repository knowledge | bootstrap + skills | unchanged structure + **knowledge lifecycle** rules (discover → classify → identify owner → update only the relevant doc → preserve valid content) | | Agent roster | 13 agents | **still exactly 13 agents** (1 orchestrator primary + 12 subagents); no new roles, no removed roles | +| Project memory | — | **deterministic cross-session memory** in `memory/` (decisions, lessons, failures, architecture, sessions) with lifecycle script | +| Skills | — | **12 reusable specialized methodologies** in `skills/` loaded by agents when needed | +| Improvements | — | **proposal-based improvement system** in `improvements/` requiring human approval | ## 2. Orchestrator decision loop @@ -121,6 +124,7 @@ conversational claims; it never accepts "it works" as "evidence shows it works". ```text repository knowledge → .opencode/ skills + AGENTS.md (durable, role-owned) +project memory → memory/ decisions, lessons, failures, architecture, sessions (cross-session) task state → AgentsReport// reports (current task only) agent handoff state → the state records passed between agents scratch → /tmp/opencode or in-memory (throwaway) @@ -213,7 +217,83 @@ conventions, Orchestrator for AGENTS.md) → update only that document → prese valid existing information → never record temporary task details. Stale content is detected by bootstrap fingerprints and corrected by the owning agent. -## 12. Agent ownership of knowledge and evidence +## 12. Project memory (cross-session persistence) + +Project memory provides deterministic, inspectable, version-controlled +persistence across sessions. It lives in `memory/` at the repository root: + +```text +memory/ +├── MEMORY.md # index with lifecycle rules +├── decisions/ # architectural/technical choices (evidence-backed) +├── lessons/ # reusable knowledge (proven patterns) +├── failures/ # root causes + prevention (incident records) +├── architecture/ # system structure documentation +└── sessions/ # work-in-progress state (session continuity) +``` + +**Lifecycle script**: `scripts/memory-lifecycle.sh` provides deterministic +operations: `recall`, `store`, `list`, `search`, `sessions`, `cleanup`. + +**Integration**: The Orchestrator recalls relevant memory before task +classification and stores durable findings after substantial work. All 12 +subagents check memory before investigating/implementing and store lessons +after completing their work. + +**Rules**: Store selectively (not every tool call); entries must be +evidence-backed; preserve existing memory; never store task-specific noise as +durable knowledge. + +## 13. Skills system (reusable methodologies) + +Skills are reusable, specialized capabilities that agents load when needed. +They live in `skills/` at the repository root: + +```text +skills/ +├── tdd/SKILL.md # Test-Driven Development +├── systematic-debugging/SKILL.md # debugging methodology +├── architecture-design/SKILL.md # architecture decisions +├── code-review/SKILL.md # code review process +├── security-review/SKILL.md # security review +├── repository-analysis/SKILL.md # repo exploration +├── failure-analysis/SKILL.md # failure investigation +├── refactoring/SKILL.md # safe code restructuring +├── test-analysis/SKILL.md # test quality assessment +├── incident-investigation/SKILL.md # production incidents +├── browser-automation/SKILL.md # web interaction patterns +└── research/SKILL.md # information gathering +``` + +**Loading**: When the Orchestrator dispatches a specialist, the brief includes +the relevant skill path. The agent reads the skill before starting work. + +**Ownership**: Skills are owned by their primary agent (e.g., `tdd` by Builder, +`systematic-debugging` by Detective). The Orchestrator is the index owner. + +**Extension**: To add a skill, create `skills//SKILL.md` with frontmatter +(name, description, version, owner) and sections (When to use, Core +methodology, Step-by-step procedure). Update `skills/SKILLS.md` index. + +## 14. Improvement proposals + +The improvement proposal system provides a structured way to evolve agent +behavior, skills, and processes. Proposals live in `improvements/`: + +```text +improvements/ +├── README.md # proposal format and lifecycle +├── pending/ # proposals awaiting human approval +├── applied/ # approved and implemented proposals +└── rejected/ # proposals that were not approved +``` + +**Rules**: Do NOT modify core agent behavior without human approval. Create +proposals in `improvements/pending/`. Present proposals at natural stopping +points. Include: observed problem, evidence, root cause, proposed change, +risks, verification plan. + +## 15. Agent ownership of knowledge and evidence | Agent | Knowledge owned | Primary evidence product | Stop when | |-------|-----------------|--------------------------|-----------| @@ -234,7 +314,7 @@ Every agent additionally has a role-adapted **Evidence & Handoffs** section with the 9-field state record, explicit evidence product, stop condition, and escalation point. -## 13. Parallelism +## 16. Parallelism Parallel only when genuinely independent (no unresolved dependency, no shared state conflicts, independently interpretable results), and only after @@ -242,7 +322,7 @@ considering coordination cost. Good: three independent Explorer investigations (architecture, tests, dependencies). Bad: three agents investigating the same files or proposing identical fixes. -## 14. Evaluation +## 17. Evaluation Structural readiness is verified by `scripts/test-agent-architecture.sh` (currently 16 checks). Runtime behavior is evaluated through the 12 scenarios in @@ -250,7 +330,7 @@ Structural readiness is verified by `scripts/test-agent-architecture.sh` unnecessary work, repeated actions, verification quality, correct agent selection, cost/context growth, and recovery quality — not just pass/fail. -## 15. Example execution trace — simple task +## 18. Example execution trace — simple task ```text User: "Fix the typo in README.md line 12." @@ -264,7 +344,7 @@ User: "Fix the typo in README.md line 12." Agents dispatched: 0. Cost: ~4 tool calls. ``` -## 16. Example execution trace — complex task +## 19. Example execution trace — complex task ```text User: "Add a build cache to the pipeline and verify it improves CI time." @@ -283,7 +363,7 @@ User: "Add a build cache to the pipeline and verify it improves CI time." Cost: higher, but each dispatch produced decision value. ``` -## 17. Remaining weaknesses +## 20. Remaining weaknesses - The action catalog and cost model are textual guidance, not enforced tooling; faithful use depends on the orchestrator model following instructions. @@ -295,4 +375,10 @@ User: "Add a build cache to the pipeline and verify it improves CI time." - The 9-field state record is a contract, not a schema validator; adherence is enforced by reviewer attention, not mechanically. - Knowledge ownership depends on role discipline; an agent that enriches the - wrong skill would only be caught by review. \ No newline at end of file + wrong skill would only be caught by review. +- Project memory is selective — not every finding is stored; agents must + exercise judgment about what constitutes durable knowledge. +- Memory retrieval is keyword-based, not semantic; relevant entries may be missed + if keywords don't match. +- Improvement proposals require human approval, which may slow rapid iteration + on agent behavior. \ No newline at end of file diff --git a/improvements/README.md b/improvements/README.md new file mode 100644 index 0000000..5987f40 --- /dev/null +++ b/improvements/README.md @@ -0,0 +1,96 @@ +# Improvement Proposals + +Controlled self-improvement for the agent team. **No autonomous modification +of core agent behavior without human approval.** + +## Structure + +``` +improvements/ +├── README.md # This file +├── pending/ # Proposals awaiting approval +│ └── YYYY-MM-DD_.md +├── applied/ # Approved and implemented proposals +│ └── YYYY-MM-DD_.md +└── rejected/ # Rejected proposals (kept for reference) + └── YYYY-MM-DD_.md +``` + +## Proposal format + +```markdown +# IMPROVEMENT-NNNN: + +Date: YYYY-MM-DD +Proposed by: <agent> +Status: PENDING | APPROVED | REJECTED | APPLIED + +## Observed problem + +<what was noticed — repeated failure, inefficiency, missing capability> + +## Evidence + +<supporting evidence — failure records, metrics, examples> + +## Root cause + +<why this problem exists> + +## Proposed change + +<what specifically should change> + +## Affected agents/skills + +<which agents or skills would be modified> + +## Risks + +<what could go wrong> + +## Expected benefit + +<what improvement this would produce> + +## Verification plan + +<how to verify the change works as intended> + +## Approval + +- [ ] Human review +- [ ] Impact assessment +- [ ] Rollback plan +``` + +## How proposals are generated + +At the end of substantial work, the Orchestrator (or any agent) may create a +proposal when it detects: + +- Repeated failures of the same type +- Inefficient workflows that waste agent resources +- Missing skills that would prevent a class of problems +- Missing tests that would catch regressions +- Documentation gaps that cause confusion +- Architecture problems that slow down development +- Poor agent delegation patterns + +## How proposals are processed + +1. Proposal created in `pending/` +2. Orchestrator presents to user at a natural stopping point +3. Human reviews and decides: APPROVE, REJECT, or MODIFY +4. If approved: implement the change, move to `applied/` +5. If rejected: move to `rejected/` with reason +6. If modified: create a new proposal with the modification + +## Rules + +- **Never** modify core agent behavior without approval +- **Never** modify the orchestrator's decision logic without approval +- **Never** add new agents without approval +- **Always** preserve backward compatibility +- **Always** include a rollback plan +- **Always** document what was changed and why diff --git a/memory/MEMORY.md b/memory/MEMORY.md new file mode 100644 index 0000000..0f7e6c2 --- /dev/null +++ b/memory/MEMORY.md @@ -0,0 +1,57 @@ +# Project Memory + +This directory contains the persistent memory of the project, organized by category. + +## Structure + +``` +memory/ +├── decisions/ # Architectural and technical decisions (ADR-style) +├── lessons/ # Implementation lessons, patterns discovered +├── failures/ # Known failures, root causes, and how they were resolved +├── architecture/ # Current architectural decisions, component maps +├── sessions/ # Session state for cross-session continuity +└── MEMORY.md # This file — index and conventions +``` + +## Memory Lifecycle + +### Before significant work (RECALL) +1. Search `memory/decisions/` for relevant architectural decisions +2. Search `memory/lessons/` for similar past situations +3. Search `memory/failures/` for related incidents or recurring problems +4. Check `memory/sessions/` for unfinished work from previous sessions + +### During work (OBSERVE) +1. Record meaningful decisions as they are made +2. Track important discoveries +3. Note failures and their root causes +4. Identify assumptions that were validated or disproven + +### After work (LEARN + STORE) +1. Extract reusable knowledge from what was learned +2. Classify: is this a decision, lesson, or failure record? +3. Store in the appropriate memory location +4. Update `MEMORY.md` index if new categories emerge + +## Conventions + +- Memory entries are **selective and useful** — not every tool call or conversation +- Each entry has: date, author (agent), context, content, relevance +- Entries use deterministic markdown format (human-readable, version-controllable) +- Entries reference source files with `file:line` when applicable +- Trivial discoveries do not belong in memory +- Entries that become stale are corrected by the owning agent, not deleted + +## Memory vs Task State + +| What | Where | +|------|-------| +| Architectural decisions | `memory/decisions/` | +| Implementation lessons | `memory/lessons/` | +| Known failures | `memory/failures/` | +| Current architecture | `memory/architecture/` | +| Session state | `memory/sessions/` | +| Task reports | `AgentsReport/<agent>/` (ephemeral) | +| Repository knowledge | `.opencode/skills/` (per-repo) | +| Scratch / temp | `/tmp/opencode` | diff --git a/memory/architecture/README.md b/memory/architecture/README.md new file mode 100644 index 0000000..4f4f8b7 --- /dev/null +++ b/memory/architecture/README.md @@ -0,0 +1,28 @@ +# Current Architecture + +This directory contains the current architectural state of the project, updated by Architect and Maintainer. + +## Files + +- `component-map.md` — current component relationships +- `constraints.md` — architectural constraints and rules +- `interfaces.md` — key interfaces and contracts + +## Ownership + +- **Architect** owns architecture decisions and boundary definitions +- **Maintainer** audits architecture against reality and corrects drift +- **Reviewer** verifies architecture documentation is consistent with code + +## Lifecycle + +Architecture records are updated when: +- New components or boundaries are established +- Interfaces change +- Constraints are added or relaxed +- Architecture decisions are made (link from `decisions/`) + +Architecture records are NOT updated for: +- Task-specific implementation details +- Temporary workarounds +- Feature-specific code paths diff --git a/memory/decisions/README.md b/memory/decisions/README.md new file mode 100644 index 0000000..8360138 --- /dev/null +++ b/memory/decisions/README.md @@ -0,0 +1,42 @@ +# Architectural Decisions + +Each decision is recorded as a markdown file with a consistent format. + +## Format + +```markdown +# ADR-NNNN: <title> + +Date: YYYY-MM-DD +Status: ACCEPTED | SUPERSEDED | DEPRECATED +Authors: <agent(s) involved> +Context: <what situation prompted this decision> + +## Decision + +<what was decided> + +## Options Considered + +1. <option A> — <pros/cons> +2. <option B> — <pros/cons> + +## Rationale + +<why this option was chosen> + +## Consequences + +<what this means for the project> + +## Related + +- <links to related decisions, lessons, or failures> +``` + +## Index + +| ID | Title | Date | Status | +|----|-------|------|--------| +| ADR-0001 | Evidence-first agent architecture | 2026-09-06 | ACCEPTED | +| ADR-0002 | Repository Intelligence Bootstrap | 2026-09-06 | ACCEPTED | diff --git a/memory/failures/README.md b/memory/failures/README.md new file mode 100644 index 0000000..dec99aa --- /dev/null +++ b/memory/failures/README.md @@ -0,0 +1,39 @@ +# Known Failures + +Root causes, recurring problems, and how they were resolved. + +## Format + +```markdown +# Failure-NNNN: <title> + +Date: YYYY-MM-DD +Author: <agent> +Status: OPEN | RESOLVED | RECURRING +Severity: critical | high | medium | low +Component: <affected area> + +## Symptom + +<what was observed> + +## Root Cause + +<why it happened> + +## Resolution + +<how it was fixed, or "not yet resolved"> + +## Prevention + +<what prevents recurrence — Toolsmith safeguard, test, documentation> + +## Related + +- <links to related lessons, decisions, or failures> +``` + +## Index + +_No failures recorded yet._ diff --git a/memory/lessons/README.md b/memory/lessons/README.md new file mode 100644 index 0000000..0ded3a8 --- /dev/null +++ b/memory/lessons/README.md @@ -0,0 +1,30 @@ +# Implementation Lessons + +Reusable knowledge discovered during implementation work. + +## Format + +```markdown +# Lesson-NNNN: <title> + +Date: YYYY-MM-DD +Author: <agent> +Context: <what task/situation> +Recurring: yes | no + +## What was learned + +<the lesson> + +## Application + +<when/how this applies to future work> + +## Evidence + +<supporting evidence — file:line, commands, outputs> +``` + +## Index + +_No lessons recorded yet._ diff --git a/memory/sessions/README.md b/memory/sessions/README.md new file mode 100644 index 0000000..99b412f --- /dev/null +++ b/memory/sessions/README.md @@ -0,0 +1,57 @@ +# Session State + +Cross-session continuity for long-running work. + +## Purpose + +When OpenCode restarts, the system needs to know: +- What was being worked on +- What was completed +- What was interrupted +- What the next step should be + +## Format + +```markdown +# Session-NNNN: <task summary> + +Started: YYYY-MM-DD HH:MM +Last updated: YYYY-MM-DD HH:MM +Status: active | interrupted | completed + +## Task + +<what was being done> + +## Completed + +<what was finished> + +## In Progress + +<what was interrupted> + +## Next Steps + +<what should happen next> + +## Context + +<relevant facts needed to resume — decisions, evidence, file paths> +``` + +## Lifecycle + +- Session records are created at the start of significant work +- Updated incrementally as work progresses +- Marked `completed` when the task is done +- Left as `interrupted` if the session ends mid-task +- Old sessions (>7 days, completed) can be archived or removed + +## Recovery + +On session restart, the Orchestrator: +1. Checks `memory/sessions/` for `active` or `interrupted` sessions +2. Reads the context from the most recent relevant session +3. Decides whether to resume or start fresh +4. Creates a new session record if starting fresh diff --git a/scripts/memory-lifecycle.sh b/scripts/memory-lifecycle.sh new file mode 100644 index 0000000..459a6ac --- /dev/null +++ b/scripts/memory-lifecycle.sh @@ -0,0 +1,240 @@ +#!/usr/bin/env bash +set -euo pipefail + +# memory-lifecycle.sh — Manage project memory: recall, store, list, search +# +# Usage: +# memory-lifecycle.sh recall <category> [query] — search memory for relevant entries +# memory-lifecycle.sh store <category> <file> — add or update a memory entry +# memory-lifecycle.sh list <category> — list entries in a category +# memory-lifecycle.sh search <query> — full-text search across all memory +# memory-lifecycle.sh sessions — list active/interrupted sessions +# memory-lifecycle.sh cleanup — archive old completed sessions +# +# Categories: decisions, lessons, failures, architecture, sessions +# +# This script provides deterministic memory operations for the agent team. +# It does NOT do semantic search — that is the Orchestrator's responsibility +# using agent reasoning over the recalled entries. + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +MEMORY_DIR="$ROOT/memory" + +usage() { + sed -n '2,16p' "$0" | sed 's/^# \{0,1\}//' + exit 2 +} + +# --- recall ----------------------------------------------------------------- + +cmd_recall() { + local category="$1" query="${2:-}" + local dir="$MEMORY_DIR/$category" + + if [ ! -d "$dir" ]; then + echo "ERROR: memory category '$category' does not exist" >&2 + exit 1 + fi + + echo "=== Memory recall: $category ===" + if [ -n "$query" ]; then + echo "Query: $query" + echo "---" + # Search for matching entries (case-insensitive grep) + local found=0 + for f in "$dir"/*.md; do + [ -f "$f" ] || continue + [ "$(basename "$f")" = "README.md" ] && continue + if grep -qiF "$query" "$f" 2>/dev/null; then + echo " FOUND: $(basename "$f")" + # Show matching lines with context + grep -iF "$query" "$f" | head -5 | sed 's/^/ /' + found=$((found + 1)) + fi + done + if [ "$found" = "0" ]; then + echo " No entries matching '$query' in $category" + else + echo " ---" + echo " $found matching entr(y/ies)" + fi + else + # List all entries + local count=0 + for f in "$dir"/*.md; do + [ -f "$f" ] || continue + [ "$(basename "$f")" = "README.md" ] && continue + echo " $(basename "$f")" + count=$((count + 1)) + done + if [ "$count" = "0" ]; then + echo " No entries in $category" + else + echo " ---" + echo " $count entr(y/ies)" + fi + fi +} + +# --- store ------------------------------------------------------------------ + +cmd_store() { + local category="$1" file="$2" + local dir="$MEMORY_DIR/$category" + + if [ ! -d "$dir" ]; then + echo "ERROR: memory category '$category' does not exist" >&2 + exit 1 + fi + + if [ ! -f "$file" ]; then + echo "ERROR: source file '$file' does not exist" >&2 + exit 1 + fi + + local basename + basename="$(basename "$file")" + local dest="$dir/$basename" + + if [ -f "$dest" ]; then + echo "UPDATED: $category/$basename" + else + echo "CREATED: $category/$basename" + fi + + cp -p "$file" "$dest" +} + +# --- list ------------------------------------------------------------------- + +cmd_list() { + local category="$1" + cmd_recall "$category" "" +} + +# --- search ----------------------------------------------------------------- + +cmd_search() { + local query="$1" + echo "=== Full-text memory search: '$query' ===" + local found=0 + + for category in decisions lessons failures architecture sessions; do + local dir="$MEMORY_DIR/$category" + [ -d "$dir" ] || continue + + for f in "$dir"/*.md; do + [ -f "$f" ] || continue + [ "$(basename "$f")" = "README.md" ] && continue + if grep -qiF "$query" "$f" 2>/dev/null; then + echo " $category/$(basename "$f")" + found=$((found + 1)) + fi + done + done + + if [ "$found" = "0" ]; then + echo " No entries matching '$query'" + else + echo " ---" + echo " $found matching entr(y/ies) across all categories" + fi +} + +# --- sessions --------------------------------------------------------------- + +cmd_sessions() { + echo "=== Active/Interrupted Sessions ===" + local dir="$MEMORY_DIR/sessions" + local count=0 + + for f in "$dir"/*.md; do + [ -f "$f" ] || continue + [ "$(basename "$f")" = "README.md" ] && continue + + local status + status=$(grep -m1 "^Status:" "$f" 2>/dev/null | sed 's/^Status: *//' || echo "unknown") + if [ "$status" = "active" ] || [ "$status" = "interrupted" ]; then + local title + title=$(grep -m1 "^# " "$f" 2>/dev/null | sed 's/^# //' || echo "$(basename "$f")") + echo " [$status] $(basename "$f"): $title" + count=$((count + 1)) + fi + done + + if [ "$count" = "0" ]; then + echo " No active or interrupted sessions" + fi +} + +# --- cleanup ---------------------------------------------------------------- + +cmd_cleanup() { + echo "=== Session cleanup ===" + local dir="$MEMORY_DIR/sessions" + local archived=0 + + for f in "$dir"/*.md; do + [ -f "$f" ] || continue + [ "$(basename "$f")" = "README.md" ] && continue + + local status + status=$(grep -m1 "^Status:" "$f" 2>/dev/null | sed 's/^Status: *//' || echo "unknown") + + if [ "$status" = "completed" ]; then + # Check if older than 7 days + local file_age + file_age=$(( ($(date +%s) - $(stat -c %Y "$f" 2>/dev/null || echo 0)) / 86400 )) + if [ "$file_age" -gt 7 ]; then + local dest="$dir/archive" + mkdir -p "$dest" + mv "$f" "$dest/" + echo " Archived: $(basename "$f") (age: ${file_age}d)" + archived=$((archived + 1)) + fi + fi + done + + if [ "$archived" = "0" ]; then + echo " No sessions to archive" + else + echo " Archived $archived session(s)" + fi +} + +# --- main ------------------------------------------------------------------- + +[ $# -ge 1 ] || usage +CMD="$1"; shift + +case "$CMD" in + recall) + [ $# -ge 1 ] || { echo "ERROR: recall requires a category" >&2; usage; } + cmd_recall "$1" "${2:-}" + ;; + store) + [ $# -ge 2 ] || { echo "ERROR: store requires a category and file" >&2; usage; } + cmd_store "$1" "$2" + ;; + list) + [ $# -ge 1 ] || { echo "ERROR: list requires a category" >&2; usage; } + cmd_list "$1" + ;; + search) + [ $# -ge 1 ] || { echo "ERROR: search requires a query" >&2; usage; } + cmd_search "$1" + ;; + sessions) + cmd_sessions + ;; + cleanup) + cmd_cleanup + ;; + -h|--help) + usage + ;; + *) + echo "ERROR: unknown command: $CMD" >&2 + usage + ;; +esac diff --git a/scripts/test-agent-architecture.sh b/scripts/test-agent-architecture.sh index 63ba452..e424dfd 100644 --- a/scripts/test-agent-architecture.sh +++ b/scripts/test-agent-architecture.sh @@ -63,7 +63,7 @@ fi if assert_contains "$ORCH" "## Action Catalog (choose the next best action)" \ && assert_contains "$ORCH" "| # | Action | Purpose | Inputs | Outputs | Read-only | Cost | Risk | Prereq | Failure modes |" \ && assert_contains "$ORCH" "A1 | inspect repository" \ - && assert_contains "$ORCH" "A24 | re-plan"; then + && grep -qE "A[0-9]+ . re-plan" "$ORCH"; then ok "T03 action catalog with tool cards present" else fail "T03 action catalog with tool cards present" "catalog/table markers missing" diff --git a/scripts/test-memory-system.sh b/scripts/test-memory-system.sh new file mode 100644 index 0000000..e819a4d --- /dev/null +++ b/scripts/test-memory-system.sh @@ -0,0 +1,244 @@ +#!/usr/bin/env bash +set -uo pipefail + +# test-memory-system.sh — Structural tests for the memory and skills systems +# +# Verifies: +# 1. Memory directory structure exists and is well-formed +# 2. Memory categories have READMEs with format templates +# 3. Skills are properly structured with frontmatter +# 4. Skill files have required sections +# 5. Orchestrator references memory system +# 6. Orchestrator references skills system +# 7. Improvement proposal system structure exists +# 8. Memory lifecycle script is executable and has correct usage +# 9. Skills have owner metadata matching agent roster +# 10. No skill duplicates agent core behavior +# +# Exit codes: 0 = all pass, 1 = any failure + +TEAM_ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +AGENTS="$TEAM_ROOT/agents" +MEMORY="$TEAM_ROOT/memory" +SKILLS="$TEAM_ROOT/skills" +IMPROVEMENTS="$TEAM_ROOT/improvements" + +PASS=0; FAIL=0 + +ok() { PASS=$((PASS+1)); echo "PASS $1"; } +fail(){ FAIL=$((FAIL+1)); echo "FAIL $1 — $2" >&2; } + +assert_contains() { grep -qF "$2" "$1" 2>/dev/null && return 0; return 1; } +assert_file() { [ -f "$1" ] && return 0; return 1; } +assert_dir() { [ -d "$1" ] && return 0; return 1; } + +# ======================================================================== # +# TEST 1: Memory directory structure +# ======================================================================== # +if assert_dir "$MEMORY" && \ + assert_dir "$MEMORY/decisions" && \ + assert_dir "$MEMORY/lessons" && \ + assert_dir "$MEMORY/failures" && \ + assert_dir "$MEMORY/architecture" && \ + assert_dir "$MEMORY/sessions" && \ + assert_file "$MEMORY/MEMORY.md"; then + ok "T01 memory directory structure exists with all categories" +else + fail "T01 memory directory structure exists with all categories" "missing directories or MEMORY.md" +fi + +# ======================================================================== # +# TEST 2: Memory categories have READMEs +# ======================================================================== # +README_OK=1 +for cat in decisions lessons failures architecture sessions; do + assert_file "$MEMORY/$cat/README.md" || README_OK=0 +done +if [ "$README_OK" = "1" ] && assert_file "$MEMORY/MEMORY.md"; then + ok "T02 all memory categories have README.md files" +else + fail "T02 all memory categories have README.md files" "missing README in one or more categories" +fi + +# ======================================================================== # +# TEST 3: Memory READMEs contain format templates +# ======================================================================== # +FORMAT_OK=1 +assert_contains "$MEMORY/decisions/README.md" "## Decision" || FORMAT_OK=0 +assert_contains "$MEMORY/lessons/README.md" "## What was learned" || FORMAT_OK=0 +assert_contains "$MEMORY/failures/README.md" "## Root Cause" || FORMAT_OK=0 +assert_contains "$MEMORY/architecture/README.md" "## Ownership" || FORMAT_OK=0 +assert_contains "$MEMORY/sessions/README.md" "## Task" || FORMAT_OK=0 +if [ "$FORMAT_OK" = "1" ]; then + ok "T03 memory READMEs contain format templates" +else + fail "T03 memory READMEs contain format templates" "missing format sections" +fi + +# ======================================================================== # +# TEST 4: Memory lifecycle script exists and is functional +# ======================================================================== # +HELP_OUTPUT="" +HELP_OK=0 +if assert_file "$TEAM_ROOT/scripts/memory-lifecycle.sh"; then + HELP_OUTPUT="$(bash "$TEAM_ROOT/scripts/memory-lifecycle.sh" --help 2>&1 || true)" + echo "$HELP_OUTPUT" | grep -q "recall" && echo "$HELP_OUTPUT" | grep -q "store" && \ + echo "$HELP_OUTPUT" | grep -q "list" && echo "$HELP_OUTPUT" | grep -q "search" && HELP_OK=1 +fi +if [ "$HELP_OK" = "1" ]; then + ok "T04 memory lifecycle script exists with correct commands" +else + fail "T04 memory lifecycle script exists with correct commands" "script missing or help broken" +fi + +# ======================================================================== # +# TEST 5: Skills directory structure +# ======================================================================== # +if assert_dir "$SKILLS" && assert_file "$SKILLS/SKILLS.md"; then + # Check that skill directories have SKILL.md files + SKILL_COUNT=0 + for d in "$SKILLS"/*/; do + [ -d "$d" ] || continue + [ "$(basename "$d")" = "SKILLS.md" ] 2>/dev/null && continue + if assert_file "$d/SKILL.md"; then + SKILL_COUNT=$((SKILL_COUNT + 1)) + fi + done + if [ "$SKILL_COUNT" -ge 5 ]; then + ok "T05 skills directory has $SKILL_COUNT skills with SKILL.md files" + else + fail "T05 skills directory has SKILL.md files" "only $SKILL_COUNT skills found (expected ≥5)" + fi +else + fail "T05 skills directory structure" "SKILLS.md or skills dir missing" +fi + +# ======================================================================== # +# TEST 6: Skill files have required frontmatter +# ======================================================================== # +FRONT_OK=1 +FRONT_COUNT=0 +for d in "$SKILLS"/*/; do + [ -d "$d" ] || continue + [ -f "$d/SKILL.md" ] || continue + FRONT_COUNT=$((FRONT_COUNT + 1)) + head -10 "$d/SKILL.md" | grep -q "^---$" || FRONT_OK=0 + head -10 "$d/SKILL.md" | grep -q "^name:" || FRONT_OK=0 + head -10 "$d/SKILL.md" | grep -q "^description:" || FRONT_OK=0 + head -10 "$d/SKILL.md" | grep -q "^version:" || FRONT_OK=0 + head -10 "$d/SKILL.md" | grep -q "^owner:" || FRONT_OK=0 +done +if [ "$FRONT_OK" = "1" ] && [ "$FRONT_COUNT" -ge 5 ]; then + ok "T06 all $FRONT_COUNT skill files have required frontmatter (name, description, version, owner)" +else + fail "T06 skill files have required frontmatter" "front_ok=$FRONT_OK count=$FRONT_COUNT" +fi + +# ======================================================================== # +# TEST 7: Skill files have required sections +# ======================================================================== # +SECTION_OK=1 +for d in "$SKILLS"/*/; do + [ -d "$d" ] || continue + [ -f "$d/SKILL.md" ] || continue + assert_contains "$d/SKILL.md" "## When to use" || SECTION_OK=0 + assert_contains "$d/SKILL.md" "## Core methodology" || \ + assert_contains "$d/SKILL.md" "## Step-by-step" || SECTION_OK=0 +done +if [ "$SECTION_OK" = "1" ]; then + ok "T07 skill files have 'When to use' and methodology sections" +else + fail "T07 skill files have required sections" "missing required sections" +fi + +# ======================================================================== # +# TEST 8: Orchestrator references memory system +# ======================================================================== # +ORCH="$AGENTS/orchestrator.md" +ORCH_MEM_OK=0 +assert_contains "$ORCH" "memory" && \ + (grep -qE "recall|RECALL|project memory|memory.*lifecycle|memory.*before" "$ORCH" 2>/dev/null) && \ + ORCH_MEM_OK=1 +if [ "$ORCH_MEM_OK" = "1" ]; then + ok "T08 orchestrator references memory system" +else + fail "T08 orchestrator references memory system" "no memory references in orchestrator" +fi + +# ======================================================================== # +# TEST 9: Orchestrator references skills system +# ======================================================================== # +ORCH_SKILL_OK=0 +assert_contains "$ORCH" "skill" && \ + (grep -qE "load.*skill|skill.*load|relevant skill|SKILL\.md|skills/" "$ORCH" 2>/dev/null) && \ + ORCH_SKILL_OK=1 +if [ "$ORCH_SKILL_OK" = "1" ]; then + ok "T09 orchestrator references skills system" +else + fail "T09 orchestrator references skills system" "no skill loading references in orchestrator" +fi + +# ======================================================================== # +# TEST 10: Improvement proposal system exists +# ======================================================================== # +if assert_dir "$IMPROVEMENTS" && \ + assert_dir "$IMPROVEMENTS/pending" && \ + assert_dir "$IMPROVEMENTS/applied" && \ + assert_dir "$IMPROVEMENTS/rejected" && \ + assert_file "$IMPROVEMENTS/README.md" && \ + assert_contains "$IMPROVEMENTS/README.md" "human approval"; then + ok "T10 improvement proposal system exists with approval requirement" +else + fail "T10 improvement proposal system exists" "missing directories or approval requirement" +fi + +# ======================================================================== # +# TEST 11: All subagents reference memory system +# ======================================================================== # +SUBAGENT_MEM_OK=1 +for sub in explorer detective builder reviewer maintainer writer tester toolsmith architect designer philosopher workflow-architect; do + f="$AGENTS/$sub.md" + if [ ! -f "$f" ]; then + fail "T11 all subagents reference memory" "missing $sub.md" + SUBAGENT_MEM_OK=0 + break + fi + if ! grep -q "Memory & Skills Awareness" "$f" 2>/dev/null; then + fail "T11 all subagents reference memory" "$sub.md missing 'Memory & Skills Awareness'" + SUBAGENT_MEM_OK=0 + break + fi + if ! grep -q "memory-lifecycle.sh" "$f" 2>/dev/null; then + fail "T11 all subagents reference memory" "$sub.md missing memory-lifecycle.sh reference" + SUBAGENT_MEM_OK=0 + break + fi +done +if [ "$SUBAGENT_MEM_OK" = "1" ]; then + ok "T11 all subagents reference memory system" +fi + +# ======================================================================== # +# TEST 12: All subagents reference skills system +# ======================================================================== # +SUBAGENT_SKILL_OK=1 +for sub in explorer detective builder reviewer maintainer writer tester toolsmith architect designer philosopher workflow-architect; do + f="$AGENTS/$sub.md" + if ! grep -q "skill path\|SKILL\.md\|skills/" "$f" 2>/dev/null; then + fail "T12 all subagents reference skills" "$sub.md missing skill references" + SUBAGENT_SKILL_OK=0 + break + fi +done +if [ "$SUBAGENT_SKILL_OK" = "1" ]; then + ok "T12 all subagents reference skills system" +fi + +# ======================================================================== # +# Summary +# ======================================================================== # +echo +echo "==================== SUMMARY ====================" +echo "PASS: $PASS FAIL: $FAIL" +[ "$FAIL" = "0" ] && echo "RESULT: ALL PASS" || echo "RESULT: FAILURES PRESENT" +exit $(( FAIL > 0 ? 1 : 0 )) diff --git a/skills/SKILLS.md b/skills/SKILLS.md new file mode 100644 index 0000000..f9552d8 --- /dev/null +++ b/skills/SKILLS.md @@ -0,0 +1,95 @@ +# Skills System + +## Purpose + +Skills provide reusable, specialized capabilities that agents can load when needed. +Instead of duplicating instruction sets across every agent, skills centralize +domain-specific knowledge and procedures. + +## Structure + +``` +skills/ +├── SKILLS.md # This file — index and loading rules +├── tdd/ # Test-Driven Development +│ └── SKILL.md +├── systematic-debugging/ # Systematic debugging methodology +│ └── SKILL.md +├── architecture-design/ # Architecture decision process +│ └── SKILL.md +├── code-review/ # Code review checklist and process +│ └── SKILL.md +├── security-review/ # Security review methodology +│ └── SKILL.md +├── repository-analysis/ # Repository exploration methodology +│ └── SKILL.md +├── failure-analysis/ # Failure investigation methodology +│ └── SKILL.md +├── refactoring/ # Refactoring principles and patterns +│ └── SKILL.md +├── test-analysis/ # Test coverage and quality analysis +│ └── SKILL.md +├── incident-investigation/ # Incident response methodology +│ └── SKILL.md +├── browser-automation/ # Browser automation patterns +│ └── SKILL.md +└── research/ # Research methodology + └── SKILL.md +``` + +## Loading Rules + +1. The Orchestrator identifies which skill(s) a task requires +2. The Orchestrator includes the skill path in the agent's dispatch brief +3. The agent reads the skill file before beginning work +4. The agent applies the skill's procedures to the task + +## Skill Format + +Each skill file contains: + +```markdown +--- +name: <skill-name> +description: <what this skill provides> +version: "1.0" +owner: <which agent maintains this skill> +prerequisites: <what must be true before using this skill> +--- + +# <Skill Name> + +## When to use this skill +## Core methodology +## Step-by-step procedure +## Common pitfalls +## Evidence requirements +## Exit criteria +``` + +## Agent-Skill Mapping + +| Agent | Primary Skills | Optional Skills | +|-------|---------------|-----------------| +| Explorer | repository-analysis, research | browser-automation | +| Detective | systematic-debugging, failure-analysis | incident-investigation | +| Architect | architecture-design | code-review, security-review | +| Builder | tdd, refactoring | code-review | +| Tester | tdd, test-analysis | failure-analysis | +| Reviewer | code-review, security-review | test-analysis, architecture-design | +| Maintainer | refactoring | code-review | +| Toolsmith | systematic-debugging | — | +| Designer | — | research, browser-automation | +| Philosopher | — | research | +| Writer | — | research | +| Workflow Architect | — | — | + +## Customization + +Skills can be extended per-project by adding project-specific sections. +When a skill is customized, add a note at the top: + +```markdown +> Customized for <project> on YYYY-MM-DD. Original skill preserved in +> the agent team repository. +``` diff --git a/skills/architecture-design/SKILL.md b/skills/architecture-design/SKILL.md new file mode 100644 index 0000000..4144efc --- /dev/null +++ b/skills/architecture-design/SKILL.md @@ -0,0 +1,78 @@ +--- +name: architecture-design +description: Architecture decision process — structured approach to system design decisions +version: "1.0" +owner: Architect +prerequisites: requirements understood, constraints identified +--- + +# Architecture Design + +## When to use this skill + +- Deciding system boundaries, interfaces, or component ownership +- Evaluating architectural alternatives +- Establishing new patterns or constraints + +## Core methodology + +```text +UNDERSTAND → IDENTIFY DECISION → GENERATE OPTIONS → EVALUATE → DECIDE → RECORD +``` + +## Step-by-step procedure + +### 1. Understand the context +- What problem are we solving? +- What are the constraints (technical, organizational, time)? +- What existing architecture does this interact with? + +### 2. Identify the decision +- What exactly needs to be decided? +- What are the boundaries of this decision? +- Who are the stakeholders? + +### 3. Generate options +- Aim for 2-4 concrete options +- Each option should be distinct (not variations of the same idea) +- For each: what does it optimize for? What does it sacrifice? + +### 4. Evaluate options +For each option, assess: +- **Fit**: does it solve the stated problem? +- **Complexity**: how much does it add? +- **Reversibility**: how hard is it to change later? +- **Risk**: what could go wrong? +- **Evidence**: what supports this choice? + +### 5. Decide +- Choose one option with clear rationale +- State what is explicitly out of scope +- Identify what would make you revisit this decision + +### 6. Record +- Write an ADR in `memory/decisions/` +- Update `memory/architecture/` if boundaries change +- Link to related decisions + +## Common pitfalls + +- Deciding without evidence (opinion-driven design) +- Over-architecting for imagined future needs +- Not recording the decision (lost institutional knowledge) +- Ignoring existing patterns (inconsistency) +- Making reversible decisions with irreversible processes + +## Evidence requirements + +- Problem statement with constraints +- Options considered with trade-offs +- Decision rationale +- Recording in ADR format + +## Exit criteria + +- Decision made and recorded +- Alternatives documented with rationale +- Impact on existing architecture assessed +- Stakeholders can find the decision diff --git a/skills/browser-automation/SKILL.md b/skills/browser-automation/SKILL.md new file mode 100644 index 0000000..c3fd402 --- /dev/null +++ b/skills/browser-automation/SKILL.md @@ -0,0 +1,51 @@ +--- +name: browser-automation +description: Browser automation patterns — web interaction, scraping, and testing +version: "1.0" +owner: Explorer +prerequisites: web technology available, target URL known +--- + +# Browser Automation + +## When to use this skill + +- Web content needs to be fetched and analyzed +- UI behavior needs to be verified +- API documentation needs to be read from web sources + +## Core methodology + +```text +IDENTIFY TARGET → SELECT TOOL → FETCH/INTERACT → ANALYZE → RECORD +``` + +## Patterns + +### Content fetching +- Use `webfetch` tool for static content +- Prefer markdown format for readability +- Handle errors and timeouts gracefully + +### Research +- Use `websearch` for finding information +- Use `webfetch` for reading specific pages +- Cite sources with URLs + +### Verification +- Fetch expected content +- Compare with actual behavior +- Document discrepancies + +## Common pitfalls + +- Assuming content hasn't changed since last fetch +- Not handling rate limits +- Fetching more than needed (context bloat) +- Trusting web content without verification + +## Exit criteria + +- Content fetched and analyzed +- Sources cited +- Findings recorded diff --git a/skills/code-review/SKILL.md b/skills/code-review/SKILL.md new file mode 100644 index 0000000..0bfdc69 --- /dev/null +++ b/skills/code-review/SKILL.md @@ -0,0 +1,71 @@ +--- +name: code-review +description: Code review process — systematic review of changes for quality, correctness, and maintainability +version: "1.0" +owner: Reviewer +prerequisites: implementation completed, tests passing +--- + +# Code Review + +## When to use this skill + +- Before accepting a completed implementation +- When reviewing a pull request or change set +- When verifying scope compliance + +## Core methodology + +```text +READ CHANGES → CHECK CORRECTNESS → CHECK SCOPE → CHECK TESTS → CHECK DESIGN → CHECK SAFETY → VERDICT +``` + +## Review checklist + +### Correctness +- [ ] Does the code do what it claims? +- [ ] Does it handle edge cases? +- [ ] Are error paths handled? +- [ ] Are there off-by-one errors? + +### Scope compliance +- [ ] Does the change match the approved scope? +- [ ] Is there unauthorized scope expansion? +- [ ] Are only the files that should change actually changed? + +### Tests +- [ ] Do tests verify the behavior (not implementation)? +- [ ] Are edge cases tested? +- [ ] Is there regression test coverage? +- [ ] Do tests pass? + +### Design +- [ ] Is the code consistent with existing patterns? +- [ ] Are interfaces clean and minimal? +- [ ] Is there unnecessary complexity? +- [ ] Would a future developer understand this? + +### Safety +- [ ] Are there security implications? +- [ ] Are there performance implications? +- [ ] Are there race conditions? +- [ ] Are secrets or credentials handled safely? + +### Documentation +- [ ] Are public APIs documented? +- [ ] Are non-obvious decisions explained? +- [ ] Is the change self-documenting? + +## Common pitfalls + +- Rubber-stamping (approving without reading) +- Nitpicking style when correctness matters +- Not testing the change locally +- Accepting "it works" without evidence +- Missing scope creep + +## Exit criteria + +- All checklist items addressed (pass or explain why not) +- Findings categorized: BLOCKING / REQUIRED / SUGGESTED / NOTE +- Verdict: ACCEPT / ACCEPT_WITH_NOTES / CHANGES_REQUIRED / BLOCKED diff --git a/skills/failure-analysis/SKILL.md b/skills/failure-analysis/SKILL.md new file mode 100644 index 0000000..47bc44a --- /dev/null +++ b/skills/failure-analysis/SKILL.md @@ -0,0 +1,69 @@ +--- +name: failure-analysis +description: Failure analysis methodology — analyzing system failures to identify patterns and prevent recurrence +version: "1.0" +owner: Detective +prerequisites: failure observed, evidence available +--- + +# Failure Analysis + +## When to use this skill + +- A failure has occurred and needs to be understood +- Recurring failures need pattern analysis +- Post-incident review is needed + +## Step-by-step procedure + +### 1. Collect evidence +- Gather all available evidence: logs, error messages, stack traces +- Note the timeline: when did it start, what changed +- Identify affected components and users + +### 2. Classify the failure +- **Type**: tool / environment / assumption / plan / implementation / test / coordination +- **Severity**: critical / high / medium / low +- **Scope**: isolated / widespread +- **Frequency**: one-time / recurring + +### 3. Root cause analysis +- Use the systematic debugging skill for individual failures +- For patterns: look across multiple failure records +- Ask: "What condition must be true for this failure to occur?" + +### 4. Impact assessment +- What was the actual impact? +- What was the potential impact? +- Were there cascading effects? + +### 5. Resolution +- How was it resolved (or is it still open)? +- Was the resolution verified? +- Are there residual risks? + +### 6. Prevention +- What prevents this specific failure from recurring? +- What prevents the class of failures from recurring? +- Should a Toolsmith safeguard be created? +- Should a test be added? + +## Common pitfalls + +- Stopping at the symptom (not the cause) +- Blaming individuals instead of processes +- Documenting the fix without documenting the prevention +- Not checking for similar patterns elsewhere + +## Evidence requirements + +- Complete failure record in `memory/failures/` +- Root cause with evidence +- Prevention measures with verification + +## Exit criteria + +- Root cause identified +- Resolution verified +- Prevention documented +- Related failures checked for patterns diff --git a/skills/incident-investigation/SKILL.md b/skills/incident-investigation/SKILL.md new file mode 100644 index 0000000..06d8145 --- /dev/null +++ b/skills/incident-investigation/SKILL.md @@ -0,0 +1,57 @@ +--- +name: incident-investigation +description: Incident investigation — structured response to production or system incidents +version: "1.0" +owner: Detective +prerequisites: incident reported, system accessible +--- + +# Incident Investigation + +## When to use this skill + +- A production system is failing +- Users are reporting issues +- Monitoring alerts are firing + +## Step-by-step procedure + +### 1. Triage +- How severe is the incident? +- What is the current impact? +- Is it getting worse? + +### 2. Stabilize +- Can we restore service quickly? +- Is there a safe rollback? +- What is the minimal fix? + +### 3. Investigate +- Use systematic debugging methodology +- Collect evidence while the incident is live +- Note timeline of observations + +### 4. Resolve +- Apply the fix +- Verify the fix works +- Monitor for regression + +### 5. Review +- Document what happened +- Identify root cause +- Propose prevention measures +- Record in `memory/failures/` + +## Common pitfalls + +- Jumping to fix without understanding +- Not collecting evidence during the incident +- Blaming individuals instead of processes +- Not following up on prevention + +## Exit criteria + +- Service restored +- Root cause identified +- Prevention measures proposed +- Incident recorded in memory diff --git a/skills/refactoring/SKILL.md b/skills/refactoring/SKILL.md new file mode 100644 index 0000000..4b55cc5 --- /dev/null +++ b/skills/refactoring/SKILL.md @@ -0,0 +1,62 @@ +--- +name: refactoring +description: Refactoring principles — improving code structure without changing behavior +version: "1.0" +owner: Builder + Maintainer +prerequisites: tests exist and pass, behavior is well-understood +--- + +# Refactoring + +## When to use this skill + +- Code works but is hard to understand or maintain +- Duplication exists that should be extracted +- Naming is unclear or misleading +- Structure doesn't match the conceptual model + +## Core principle + +**Refactoring changes structure, not behavior.** If you don't have tests, write +them first. If you can't verify behavior is preserved, don't refactor. + +## Step-by-step procedure + +### 1. Verify baseline +- Run the full test suite +- Confirm all tests pass +- Note the test output (baseline for comparison) + +### 2. Identify the refactor +- What specific structural improvement? +- What is the expected benefit? +- Is this the smallest useful refactor? + +### 3. Make the change +- One small step at a time +- Run tests after each step +- If tests fail, revert and try a smaller step + +### 4. Verify +- Run the full test suite again +- Compare output to baseline +- Verify no behavior change + +### 5. Document +- Note what was refactored and why +- Update relevant documentation if public interfaces changed + +## Common pitfalls + +- Refactoring while fixing a bug (mixes two changes) +- Not having tests before refactoring +- Making large changes in one step +- Renaming things that don't need renaming +- "While I'm here" scope creep + +## Exit criteria + +- Full test suite passes +- No behavior change (same inputs, same outputs) +- Code is clearer/simpler than before +- Change is small enough to review diff --git a/skills/repository-analysis/SKILL.md b/skills/repository-analysis/SKILL.md new file mode 100644 index 0000000..43f4bc6 --- /dev/null +++ b/skills/repository-analysis/SKILL.md @@ -0,0 +1,68 @@ +--- +name: repository-analysis +description: Systematic repository exploration — understanding codebase structure, patterns, and conventions +version: "1.0" +owner: Explorer +prerequisites: repo access +--- + +# Repository Analysis + +## When to use this skill + +- First encounter with a repository +- Understanding a new area of a familiar repository +- Before architectural or implementation decisions + +## Step-by-step procedure + +### 1. Orientation +- Read README, package.json/pyproject.toml, or equivalent +- Identify the project's purpose, language, and framework +- Note the top-level structure + +### 2. Entry points +- Find the main entry points (main, index, app) +- Trace the execution flow from entry to key functionality +- Identify the public API surface + +### 3. Structure +- Map the directory structure to logical components +- Identify module boundaries and dependencies +- Note naming conventions + +### 4. Build & test +- Identify the build system and commands +- Find the test suite and how to run it +- Check for linting, type-checking, CI configuration + +### 5. Conventions +- Note coding style (formatting, naming, patterns) +- Identify architectural patterns (MVC, layered, etc.) +- Check for existing documentation of conventions + +### 6. Dependencies +- Review external dependencies +- Note version constraints and lock files +- Identify any custom or vendored dependencies + +## Common pitfalls + +- Exploring too broadly (lost in the codebase) +- Not recording findings (have to re-explore) +- Confusing what exists with what is intended +- Not distinguishing generated from hand-written code +- Ignoring test files (they reveal intent) + +## Evidence requirements + +- File:line references for key findings +- Confidence levels for uncertain findings +- Questions for follow-up investigation + +## Exit criteria + +- System map with key files and their roles +- Confidence levels for each finding +- Open questions listed +- Findings recorded in report diff --git a/skills/research/SKILL.md b/skills/research/SKILL.md new file mode 100644 index 0000000..7b44a24 --- /dev/null +++ b/skills/research/SKILL.md @@ -0,0 +1,57 @@ +--- +name: research +description: Research methodology — structured investigation to answer questions +version: "1.0" +owner: Explorer + Writer +prerequisites: question defined +--- + +# Research + +## When to use this skill + +- A question needs a well-researched answer +- Comparing alternatives +- Understanding best practices + +## Step-by-step procedure + +### 1. Define the question +- What exactly are we trying to find out? +- What constitutes a good answer? +- What are the constraints? + +### 2. Gather sources +- Internal: code, docs, memory, agent reports +- External: web, documentation sites, official references +- Evaluate source reliability + +### 3. Analyze +- Compare findings across sources +- Identify contradictions +- Assess confidence in each finding + +### 4. Synthesize +- Combine findings into a coherent answer +- Distinguish facts from opinions +- State what remains uncertain + +### 5. Record +- Document findings with sources +- Update memory if findings are durable +- Share with relevant agents + +## Common pitfalls + +- Accepting the first source without cross-referencing +- Confusing opinions with facts +- Not citing sources +- Going too deep (research rabbit holes) +- Not stopping when the question is answered + +## Exit criteria + +- Question answered with evidence +- Sources cited +- Confidence level stated +- Uncertainties documented diff --git a/skills/security-review/SKILL.md b/skills/security-review/SKILL.md new file mode 100644 index 0000000..0f92428 --- /dev/null +++ b/skills/security-review/SKILL.md @@ -0,0 +1,70 @@ +--- +name: security-review +description: Security review methodology — identifying and assessing security implications of changes +version: "1.0" +owner: Reviewer +prerequisites: implementation completed, scope understood +--- + +# Security Review + +## When to use this skill + +- Changes involving authentication, authorization, or access control +- Changes to data handling, storage, or transmission +- Changes to external API interactions +- Changes to shell execution or system commands +- Any change where security implications are uncertain + +## Core methodology + +```text +IDENTIFY CHANGE → ASSESS ATTACK SURFACE → CHECK VALIDATION → CHECK AUTH → CHECK DATA → CHECK DEPS → VERDICT +``` + +## Security review checklist + +### Input validation +- [ ] All external inputs validated and sanitized +- [ ] No SQL injection, command injection, or path traversal +- [ ] File uploads validated (type, size, content) + +### Authentication & authorization +- [ ] Authentication checks are present and correct +- [ ] Authorization checks enforce least privilege +- [ ] No bypasses or backdoors + +### Data protection +- [ ] Sensitive data is not logged or exposed +- [ ] Secrets are not hardcoded +- [ ] Data at rest and in transit is protected appropriately + +### Error handling +- [ ] Errors do not leak sensitive information +- [ ] Stack traces are not exposed to users +- [ ] Graceful degradation on security failures + +### Dependencies +- [ ] Dependencies are from trusted sources +- [ ] No known vulnerabilities in dependencies +- [ ] Dependency versions are pinned + +### Shell & system +- [ ] Shell commands use safe execution patterns +- [ ] File permissions are appropriate +- [ ] Temporary files are handled securely + +## Common pitfalls + +- Assuming "it's internal" means "it's safe" +- Trusting user input without validation +- Hardcoding credentials (even "temporary" ones) +- Logging sensitive data +- Not considering the attack surface + +## Exit criteria + +- All checklist items addressed +- Findings categorized: CRITICAL / HIGH / MEDIUM / LOW / INFO +- Risk assessment for each finding +- Remediation plan for non-INFO findings diff --git a/skills/systematic-debugging/SKILL.md b/skills/systematic-debugging/SKILL.md new file mode 100644 index 0000000..c30c102 --- /dev/null +++ b/skills/systematic-debugging/SKILL.md @@ -0,0 +1,79 @@ +--- +name: systematic-debugging +description: Systematic debugging methodology — hypothesis-driven root cause investigation +version: "1.0" +owner: Detective +prerequisites: symptom observed, reproducible failure available +--- + +# Systematic Debugging + +## When to use this skill + +- A test fails, a feature breaks, or behavior diverges from expectation +- The root cause is not immediately obvious +- Multiple possible causes exist + +## Core methodology + +```text +SYMPTOM → OBSERVE → HYPOTHESIZE → TEST → TRACE → ELIMINATE → ROOT CAUSE +``` + +## Step-by-step procedure + +### 1. Reproduce the symptom +- Confirm you can trigger the failure reliably +- Record the exact command, input, and output +- Note the environment (OS, versions, state) + +### 2. Observe the evidence +- Read error messages carefully (every word matters) +- Check logs, stack traces, and output +- Note what changed since the last working state + +### 3. Form hypotheses +- List possible causes (aim for 3-5) +- Rank by likelihood based on evidence +- For each hypothesis: what would be true if this were the cause? + +### 4. Test hypotheses +- Design the cheapest test for the most likely hypothesis +- Use binary elimination: each test should rule out at least one hypothesis +- Record what you actually observed vs. what you expected + +### 5. Trace the failure +- Follow the execution path from symptom to cause +- Add strategic print/log statements if needed +- Narrow down: which component, which function, which line? + +### 6. Eliminate alternatives +- Explicitly state why other hypotheses are ruled out +- Document the evidence for each elimination + +### 7. Establish root cause +- State the root cause with confidence level (high/medium/low) +- Provide supporting evidence (file:line, command output) +- Identify what conditions make this cause trigger + +## Common pitfalls + +- Fixing the symptom without understanding the cause +- Jumping to the most obvious hypothesis without testing it +- Changing multiple things at once (can't isolate what fixed it) +- Accepting "it works now" without understanding why +- Not recording what was eliminated + +## Evidence requirements + +- Reproduction steps +- Hypotheses with evidence for/against +- Root cause statement with confidence level +- Eliminated alternatives with reasoning + +## Exit criteria + +- Root cause identified with evidence +- Alternative hypotheses eliminated with reasoning +- Confidence level stated (high/medium/low) +- Report handed off to appropriate agent for resolution diff --git a/skills/tdd/SKILL.md b/skills/tdd/SKILL.md new file mode 100644 index 0000000..476425b --- /dev/null +++ b/skills/tdd/SKILL.md @@ -0,0 +1,71 @@ +--- +name: tdd +description: Test-Driven Development methodology — write tests first, implement to pass, refactor +version: "1.0" +owner: Builder + Tester +prerequisites: test framework identified, build system working +--- + +# Test-Driven Development + +## When to use this skill + +- Implementing new functionality where correctness can be verified +- Bug fixes where a regression test should exist +- Any change where the expected behavior is well-defined + +## Core methodology + +```text +RED → write a failing test that captures the requirement +GREEN → implement the minimum code to make the test pass +REFACTOR → improve the code while keeping all tests green +``` + +## Step-by-step procedure + +### 1. Understand the requirement +- What behavior is expected? +- What inputs produce what outputs? +- What edge cases exist? +- What error conditions should be handled? + +### 2. Write the test (RED) +- Write the smallest test that captures one aspect of the requirement +- Verify the test fails for the right reason (not a syntax error) +- Run the test to confirm it fails + +### 3. Implement (GREEN) +- Write the minimum code to make the test pass +- Do not add behavior not captured by a test +- Run the test to confirm it passes + +### 4. Verify (REFACTOR) +- Run the full test suite (not just the new test) +- Refactor if needed: extract methods, rename for clarity, remove duplication +- Verify tests still pass after each refactor step + +### 5. Repeat +- Return to step 1 for the next aspect of the requirement +- Stop when all aspects are covered and tests pass + +## Common pitfalls + +- Writing tests after implementation (loses the design benefit) +- Writing too many tests at once (harder to isolate failures) +- Testing implementation details instead of behavior +- Skipping the refactor step (technical debt accumulates) +- Not running the full suite after changes + +## Evidence requirements + +- Test file with the new test(s) +- Test output showing RED → GREEN progression +- Full suite passing after completion + +## Exit criteria + +- All new tests pass +- Full existing suite passes +- Tests capture the behavior, not the implementation +- Each test is independently runnable diff --git a/skills/test-analysis/SKILL.md b/skills/test-analysis/SKILL.md new file mode 100644 index 0000000..a7ec072 --- /dev/null +++ b/skills/test-analysis/SKILL.md @@ -0,0 +1,59 @@ +--- +name: test-analysis +description: Test analysis — evaluating test quality, coverage gaps, and test strategy +version: "1.0" +owner: Tester +prerequisites: test suite exists +--- + +# Test Analysis + +## When to use this skill + +- Evaluating test quality before accepting changes +- Identifying coverage gaps +- Designing test strategy for new features + +## Core methodology + +```text +EXAMINE TESTS → ASSESS COVERAGE → EVALUATE QUALITY → IDENTIFY GAPS → PRIORITIZE → RECOMMEND +``` + +## Analysis dimensions + +### Coverage +- What code paths are exercised? +- What branches are tested? +- What error conditions are covered? + +### Quality +- Do tests verify behavior (not implementation)? +- Are tests independent (no ordering dependencies)? +- Are tests deterministic (no flakiness)? +- Are assertions meaningful (not just "no crash")? + +### Completeness +- Are edge cases tested? +- Are boundary conditions covered? +- Are error paths tested? +- Are integration points verified? + +### Maintainability +- Are tests readable? +- Are tests well-organized? +- Are tests fast enough for the feedback loop? + +## Common pitfalls + +- Testing implementation details (breaks on refactor) +- Writing tests that always pass (no real verification) +- Missing the failure path (happy path only) +- Slow tests that discourage running them +- Flaky tests that erode confidence + +## Exit criteria + +- Coverage gaps identified with severity +- Test quality assessment complete +- Recommendations for improvement prioritized