Add adaptive, evidence-driven architecture upgrade: decision-loop orchestrator, action catalog, evidence-state handoffs, process quality, stop conditions

- Orchestrator: linear flow replaced by adaptive decision loop
  (UNDERSTAND→ESTIMATE→LOAD CONTEXT→CHOOSE ACTION→EXECUTE→VERIFY→RE-PLAN/STOP→LEARN)
- New sections: Task Complexity Estimation (ESTIMATE→EXECUTE→EXPAND),
  Action Catalog with 24 tool-card actions, Evidence-First State and Handoff
  Discipline (9-field state records, state separation, long-horizon persistence),
  Process Quality anti-patterns, Cost and Token Awareness, Stop Conditions,
  adaptive planning + failure recovery (7 failure types), lightweight quality gates
- All 13 agents preserved; 12 subagents gain role-adapted Evidence & Handoffs
  sections (evidence product, stop, escalation) after Repository Intelligence
- Repository Intelligence Bootstrap extended with Knowledge Lifecycle rules;
  .opencode structure unchanged (knowledge/state dirs intentionally not added)
- New docs: AGENT_ARCHITECTURE.md (full architecture incl. simple/complex
  execution traces, ownership table, weaknesses) and EVALUATION_SCENARIOS.md
  (12 runtime scenarios + scoring rubric)
- New test: scripts/test-agent-architecture.sh (16 structural checks)
- Validation: 16/16 architecture tests PASS, 11/11 bootstrap tests PASS,
  install 13/13 + permission verifier PASS, live config synced byte-identical
This commit is contained in:
Your Name
2026-09-07 12:53:30 -04:00
parent c1f5f939ad
commit fb9d91e510
17 changed files with 1277 additions and 33 deletions
+20
View File
@@ -66,6 +66,26 @@ domain, but always verify root cause against the actual source. Do not modify
- **Consume**: repo-context, build-and-test (when debugging test/build failures),
architecture (when tracing across modules)
## Evidence & Handoffs
Produce structured state records for hypotheses, failures, and handoffs — not for every diagnostic command:
```text
goal: <the symptom you are explaining>
hypothesis: <what you believe is true>
evidence: <what was observed — commands, outputs, logs, file:line>
actions_taken: <what was actually done>
result: <what happened>
verification: <how the result was confirmed — reproduction, elimination>
confidence: high | medium | low
remaining_unknowns: <what is still not known>
recommended_next_action: <what should happen next, and who owns it>
```
Your primary evidence distinguishes facts from guesses: each hypothesis must name the test that probes it and the observed result. Report confidence for the root cause AND for eliminated alternatives.
Stop when your deliverable is complete and verified per your Completion Rule; escalate when required evidence is missing or the symptom is out of scope.
## Investigation Boundary
Your job is diagnosis, and your sandbox permissions are writable. Use that only where this prompt permits: