Add adaptive, evidence-driven architecture upgrade: decision-loop orchestrator, action catalog, evidence-state handoffs, process quality, stop conditions

- Orchestrator: linear flow replaced by adaptive decision loop
  (UNDERSTAND→ESTIMATE→LOAD CONTEXT→CHOOSE ACTION→EXECUTE→VERIFY→RE-PLAN/STOP→LEARN)
- New sections: Task Complexity Estimation (ESTIMATE→EXECUTE→EXPAND),
  Action Catalog with 24 tool-card actions, Evidence-First State and Handoff
  Discipline (9-field state records, state separation, long-horizon persistence),
  Process Quality anti-patterns, Cost and Token Awareness, Stop Conditions,
  adaptive planning + failure recovery (7 failure types), lightweight quality gates
- All 13 agents preserved; 12 subagents gain role-adapted Evidence & Handoffs
  sections (evidence product, stop, escalation) after Repository Intelligence
- Repository Intelligence Bootstrap extended with Knowledge Lifecycle rules;
  .opencode structure unchanged (knowledge/state dirs intentionally not added)
- New docs: AGENT_ARCHITECTURE.md (full architecture incl. simple/complex
  execution traces, ownership table, weaknesses) and EVALUATION_SCENARIOS.md
  (12 runtime scenarios + scoring rubric)
- New test: scripts/test-agent-architecture.sh (16 structural checks)
- Validation: 16/16 architecture tests PASS, 11/11 bootstrap tests PASS,
  install 13/13 + permission verifier PASS, live config synced byte-identical
This commit is contained in:
Your Name
2026-09-07 12:53:30 -04:00
parent c1f5f939ad
commit fb9d91e510
17 changed files with 1277 additions and 33 deletions
+20
View File
@@ -63,6 +63,26 @@ validation step), add them to build-and-test and strip the
- **Owned**: `.opencode/skills/build-and-test/SKILL.md` (with Builder)
- **Consume**: conventions, repo-context, architecture (when relevant)
## Evidence & Handoffs
Produce structured state records for test results, coverage decisions, and handoffs — not for every assertion:
```text
goal: <the behavior you were asked to verify>
hypothesis: <the behavior you expect the system to exhibit> (when relevant)
evidence: <what was observed — tests run, pass/fail counts, outputs, logs>
actions_taken: <what was actually done>
result: <what happened>
verification: <the run command and its outcome>
confidence: high | medium | low
remaining_unknowns: <what is still not known — untested paths, flaky cases>
recommended_next_action: <what should happen next, and who owns it>
```
Your evidence is verification: tests run, pass/fail totals, reproduction commands, and defects by severity. Report coverage gaps honestly — never mark a requirement verified without a run.
Stop when your deliverable is complete and verified per your Completion Rule; escalate when the verification matrix lacks the facts needed to test the behavior.
## Core Philosophy
Mirror disciplined practical testing: