📖 Blueprints & Methodologies

Pentest methodology documents, operational blueprints, and security references.

← Back to list agentic-orchestrated-blueprint.md
agentic-orchestrated-blueprint.md
Download Raw
# ptSlick — Agentic Orchestrated Blueprint v1.0
**Classification**: Internal — SlickLab Agent Operations
**Build**: 2026-10-03 | **Next Review**: 2026-11-03
**Author**: Colnex (SlickLab Operational Intelligence)

---

## 0. Executive Summary

This blueprint defines how to **orchestrate multiple AI agents into coordinated, auditable, security-hardened workflows** for red team operations and security research. It covers 8 orchestration patterns, 2 communication protocols, 3 memory tiers, and the security architecture that binds them together — informed by the Paperclip CVE, MCP STDIO injection, and the 2026 agent security landscape.

**Key insight from research (2026 consensus):** Start with a single agent. Only add multi-agent orchestration when you hit a specific limitation — tool overload (15-20+ tools), context window overflow, distinct security boundaries, or genuine need for specialization. Multi-agent multiplies token cost roughly **4-15x** over single-agent. The patterns in this blueprint each pay that tax differently.

---

## 1. Architecture Principles

### 1.1 The Hierarchy of Orchestration Complexity

```
  Cost ←───────────────────────────────────────→ Complexity
                                                      
  SINGLE AGENT                                   LEVEL 0
  └─ One agent, one context, N tools                  
                                                      
  SEQUENTIAL PIPELINE                             LEVEL 1
  └─ Agent A → Agent B → Agent C                      
     Deterministic, fixed order, every step depends on previous
                                                      
  SUPERVISOR / ORCHESTRATOR-WORKER                LEVEL 2
  └─ Lead agent decomposes → delegates → synthesizes    
     Dynamic subtasks, specialist workers, single accountability
                                                      
  CONCURRENT / FAN-OUT/FAN-IN                     LEVEL 2
  └─ Same input → N parallel agents → aggregated      
     Time win: wall clock = max(agents), not sum
                                                      
  HANDOFF                                         LEVEL 3
  └─ Task passes agent-to-agent, one active at a time  
     Cheapest multi-agent pattern (no parallel context)
                                                      
  GROUP CHAT / DEBATE                             LEVEL 3
  └─ Shared thread, chat manager controls turn order   
     Best for maker-checker quality loops, capped at 3
                                                      
  MAGENTIC / ADAPTIVE PLANNING                    LEVEL 4
  └─ Manager builds live task ledger, iterates         
     For problems with no predetermined solution path
                                                      
  BLACKBOARD                                      LEVEL 4
  └─ Shared state store, agents react to events       
     Event-driven, needs conflict resolution
                                                      
  SWARM                                           LEVEL 5
  └─ Peer-to-peer routing, no central coordinator      
     Highest complexity, hardest to debug
```

**Rule:** Default to LEVEL 0. Only escalate when you can articulate the specific limitation that forces it.

### 1.2 When to Orchestrate (and When Not To)

```
QUESTION CHAIN:
┌──────────────────────────────────────────────────────┐
│ Can one agent handle this in one context window?     │
│   YES  → Single agent. Done.                         │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Is the task decomposable into fixed stages?          │
│   YES  → Sequential pipeline (Level 1)               │
│         Cost = N agents × 1 pass                      │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Are the subtasks known but independent?              │
│   YES  → Supervisor/orchestrator-worker (Level 2)    │
│         Cost = 1 planner + N workers + 1 synthesizer │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Do you need multiple perspectives on one input?      │
│   YES  → Concurrent/fan-out (Level 2)                │
│         Cost = N agents + 1 aggregator                │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Is the optimal routing unknown until runtime?        │
│   YES  → Handoff (Level 3)                           │
│         Cost = 1 active agent × depth                 │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Do you need quality verification / reflection?       │
│   YES  → Group chat / debate (Level 3)               │
│         Cost = N agents × rounds, cap at 3 agents    │
│   NO   ↓                                             │
├──────────────────────────────────────────────────────┤
│ Is the solution path unknowable upfront?             │
│   YES  → Magentic / adaptive planning (Level 4)      │
│         Cost = unbounded until plan converges        │
│   NO   → You probably don't need multi-agent         │
└──────────────────────────────────────────────────────┘
```

### 1.3 Critical Cost Data (2026 Benchmarks)

```
┌────────────────────────────────┬──────────┬─────────────┐
│ Pattern                        │ Token    │ Coordination│
│                                │ Multiple │ Overhead per │
│                                │          │ step         │
├────────────────────────────────┼──────────┼─────────────┤
│ Single agent with tools        │   ~4x    │       0ms    │
│ Sequential pipeline (4 agents) │  ~29k    │    ~950ms    │
│ Supervisor-worker (5 workers)  │  ~15x    │  ~2-3s       │
│ Concurrent (4 agents)          │   ~4x    │  ~500ms      │
│ Group chat (3 agents, 3 rnds)  │  ~9x     │  ~4-8s       │
│ Magentic (unbounded)           │  ~20x+   │ variable     │
│ Swarm (N agents × rounds)      │  N×R×2  │  N(N-1)/2    │
└────────────────────────────────┴──────────┴─────────────┘

Source: Microsoft Azure Architecture Center, Anthropic engineering data, 2026.
Single-agent-with-tools baseline: ~4x standard chat. Multi-agent starts at ~15x.
```

---

## 2. Communication Protocols

### 2.1 Protocol Topology

```
                    ┌─────────────────────┐
                    │    USER/OPERATOR    │
                    └──────────┬──────────┘
                               │
                    ┌──────────▼──────────┐
                    │  ORCHESTRATOR       │
                    │  (Agent Manager)    │
                    └──┬────────────┬─────┘
                       │            │
              ┌────────▼──┐   ┌────▼────────┐
              │ MCP CLIENT│   │ A2A CLIENT  │
              │ (tools)   │   │ (peers)     │
              └───┬───────┘   └────┬────────┘
                  │                │
         ┌────────▼────────┐  ┌───▼──────────────┐
         │ MCP Servers     │  │ Peer Agents       │
         │ ┌────────────┐  │  │ ┌────────────┐   │
         │ │ Search     │  │  │ │ Code Agent │   │
         │ │ DB Query   │  │  │ │ Research   │   │
         │ │ File Sys   │  │  │ │ Deploy     │   │
         │ │ API Gate   │  │  │ │ Analyst    │   │
         │ │ Sandbox    │  │  │ └────────────┘   │
         │ └────────────┘  │  └──────────────────┘
         └─────────────────┘

MCP = Model Context Protocol (tool access, client-server)
A2A = Agent-to-Agent protocol (peer delegation, HTTP/JSON-RPC)
```

### 2.2 Model Context Protocol (MCP) Architecture

MCP is the **USB interface for AI agents** — it standardizes how agents connect to tools, data sources, and services.

```
┌─────────────────────────────────────────────────────────┐
│ MCP Host (Orchestrator / Agent Runtime)                 │
│  ┌──────────────────────────────────────────────────┐  │
│  │ MCP Client:  Agent ↔ Server communication        │  │
│  │  - tools/call     (execute tool by name)          │  │
│  │  - resources/read (access data by URI)            │  │
│  │  - prompts/get    (retrieve prompt templates)     │  │
│  │  - session management (stateful exchanges)        │  │
│  └──────────────────────────────────────────────────┘  │
│              │                    │                     │
│     ┌────────▼────────┐   ┌──────▼───────────────┐     │
│     │ MCP Server A    │   │ MCP Server B          │     │
│     │ Search Tools    │   │ Database Tools        │     │
│     │ ─────────────── │   │ ───────────────────── │     │
│     │ web_search()    │   │ query_db()            │     │
│     │ web_extract()   │   │ read_table()          │     │
│     │ url_scan()      │   │ write_row()           │     │
│     └─────────────────┘   └───────────────────────┘     │
└─────────────────────────────────────────────────────────┘
```

**MCP Transport Types (critical security dimension):**

| Transport | How It Works | Security Risk |
|-----------|-------------|---------------|
| **STDIO** | Server runs as subprocess, communicates over stdin/stdout | **HIGH** — CVE-2026-30623, CVE-2026-30615. STDIO server definition is executable content. 200k+ instances affected, 150M+ package downloads. Tool descriptions parsed as code. |
| **HTTP/SSE** | Server listens on HTTP port, agent sends JSON-RPC | **MEDIUM** — Network exposure, needs auth. OAuth 2.1 adopted 2026. |
| **WebSocket** | Persistent bidirectional channel | **MEDIUM** — Same as HTTP + session management |

**MCP Security Hardening (from Paperclip + STDIO lessons):**

```
┌────────────────┬────────────────────────────────────────────────┐
│ Rule           │ Implementation                                 │
├────────────────┼────────────────────────────────────────────────┤
│ 1. Config is   │ Every MCP server definition is executable.     │
│    code        │ Treat tool descriptions, parameter schemas,    │
│                │ and STDIO command strings as code. Audit them. │
├────────────────┼────────────────────────────────────────────────┤
│ 2. Narrow      │ One MCP server per capability. Never a "super  │
│    servers     │ server" with every tool. Scope = blast radius. │
├────────────────┼────────────────────────────────────────────────┤
│ 3. Least       │ OAuth 2.1 + PKCE. No long-lived tokens.       │
│    privilege   │ Scopes narrow per resource. Token rotation.    │
├────────────────┼────────────────────────────────────────────────┤
│ 4. Gateway     │ Route all MCP traffic through gateway. Central │
│    pattern     │ authN/Z, policy-as-code (OPA), rate limits,    │
│                │ audit. Kill switch per server.                 │
├────────────────┼────────────────────────────────────────────────┤
│ 5. No STDIO    │ Default to HTTP/SSE or WebSocket transport.    │
│    by default  │ STDIO only when server on same host, non-root, │
│                │ read-only filesystem, minimal container.       │
├────────────────┼────────────────────────────────────────────────┤
│ 6. Supply      │ Pin server versions. Verify build provenance   │
│    chain       │ (Sigstore). Watch for ownership changes.       │
│    security    │ Treat MCP servers like npm/pip dependencies.   │
├────────────────┼────────────────────────────────────────────────┤
│ 7. Shadow MCP  │ Register all MCP servers. Inventory them.      │
│    detection   │ Unregistered servers = ungoverned = risk.      │
└────────────────┴────────────────────────────────────────────────┘
```

### 2.3 Agent-to-Agent Protocol (A2A)

A2A governs peer-level agent communication — delegation, negotiation, capability discovery.

```
┌────────────────────────────────────────────────────────────┐
│ A2A Communication Flow                                     │
│                                                            │
│  Agent A (Requester)           Agent B (Responder)         │
│  ┌──────────────────┐          ┌──────────────────┐       │
│  │ A2A Agent Card   │ ───────► │ A2A Agent Card   │       │
│  │ - capabilities   │   GET    │ - capabilities   │       │
│  │ - auth methods   │          │ - auth methods   │       │
│  │ - endpoint URL   │          │ - endpoint URL   │       │
│  └───────┬──────────┘          └───────┬──────────┘       │
│          │                             │                    │
│          │   POST /task (JSON-RPC)     │                    │
│          ├────────────────────────────►│                    │
│          │   {                         │                    │
│          │     "jsonrpc": "2.0",       │                    │
│          │     "method": "tasks/send", │                    │
│          │     "params": {             │                    │
│          │       "id": "task-001",     │                    │
│          │       "message": {          │                    │
│          │         "role": "user",     │                    │
│          │         "parts": [...]      │                    │
│          │       }                     │                    │
│          │     }                       │                    │
│          │   }                         │                    │
│          │◄────────────────────────────┤                    │
│          │   { "id": "task-001",       │                    │
│          │     "status": "accepted" }  │                    │
│          │                             │                    │
│          │   (Agent B processes...)    │                    │
│          │                             │                    │
│          │◄────────────────────────────┤                    │
│          │   POST /task (update)       │                    │
│          │   { "status": "completed",  │                    │
│          │     "artifact": {...} }     │                    │
│          └─────────────────────────────┘                    │
└────────────────────────────────────────────────────────────┘

Security properties:
  - Cryptographic signing at every hop
  - Delegation scope MUST narrow (never expand)
  - Identity attestation via AIP (Agent Identity Protocol)
  - Every A2A exchange logged to orchestration state
```

### 2.4 MCP + A2A: The Dual Foundation

```
┌────────────────────────────────────────────────────┐
│              AGENT ORCHESTRATION LAYER              │
│  ┌──────────┐  ┌──────────────┐  ┌──────────────┐ │
│  │ Planner  │  │ State/Known │  │ Quality/Ops  │ │
│  │ (delegate│  │ (memory)    │  │ (monitor)    │ │
│  │ tasks)   │  │             │  │              │ │
│  └────┬─────┘  └──────┬───────┘  └──────┬───────┘ │
│       │               │                  │          │
│  ┌────▼───────────────▼──────────────────▼───────┐ │
│  │         AGENT COMMUNICATION LAYER             │ │
│  │  ┌─────────────────┐  ┌──────────────────┐   │ │
│  │  │  MCP (tools)    │  │  A2A (peers)     │   │ │
│  │  │  client-server  │  │  peer-to-peer    │   │ │
│  │  │  tool calls     │  │  task delegation │   │ │
│  │  │  data access    │  │  capability disc │   │ │
│  │  └─────────────────┘  └──────────────────┘   │ │
│  └───────────────────────────────────────────────┘ │
└────────────────────────────────────────────────────┘

MCP answers: "How does this agent get tools?"
A2A answers: "How does this agent talk to other agents?"
```

---

## 3. Orchestration Patterns Catalog

### 3.1 Sequential Pipeline (Level 1)

```
┌──────────┐     ┌──────────┐     ┌──────────┐     ┌──────────┐
│ Agent A  │ ──► │ Agent B  │ ──► │ Agent C  │ ──► │ Agent D  │
│ Parse    │     │ Analyze  │     │ Generate │     │ Validate │
└──────────┘     └──────────┘     └──────────┘     └──────────┘
     │                │                │                │
     └─────── Shared State ────────────┴────────────────┘

When to use:     Fixed sequence, clear dependencies, batch processing
Cost profile:    N agents, each sees previous output
Failure mode:    Error cascades — bad output in step 1 poisons everything
                 No backtracking. Mitigation: checkpoint per step
Debugging:       Trivial — failure at step N means error in step N
```

**Red team application:** Recon pipeline — SpiderFoot → Amass → Nmap → Vuln scan.

### 3.2 Supervisor / Orchestrator-Worker (Level 2)

```
                         ┌──────────────┐
                         │ SUPERVISOR   │
                         │ Plans,       │
                         │ delegates,   │
                         │ synthesizes  │
                         └──┬───────┬───┘
                     ┌──────┤       ├──────┐
                     │      │       │      │
              ┌──────▼──┐┌──▼────┐┌─▼─────┐
              │ Worker 1││Worker2││Worker3│
              │ Recon   ││Exfil  ││Report │
              └─────────┘└───────┘└───────┘

When to use:     Task decomposes cleanly, specialist agents
                 Need single point of accountability
Cost profile:    1 planner + N workers + 1 synthesizer
                 3-5 workers ideal. Beyond that, context overflow.
Failure mode:    Supervisor is single point of failure.
                 Misclassification → wrong worker gets task.
Mitigation:      Supervisor uses capable model, workers use cheaper ones.
                 Cost 40-60% less than running all on capable models.
```

**Red team application:** Campaign planning — supervisor defines TTPs, workers execute in parallel (phish, exploit, recon, pivot), synthesizer builds report.

### 3.3 Concurrent / Fan-Out / Fan-In (Level 2)

```
                    ┌──────────────┐
                    │ DISPATCHER   │
                    │ (same input) │
                    └──┬───┬───┬───┘
                       │   │   │
              ┌────────▼┐ ┌▼──┐ ┌▼────────┐
              │ Agent A │ │ B │ │ Agent C │
              │ Web     │ │API│ │ DB      │
              └─────────┘ └───┘ └─────────┘
                       │   │   │
                    ┌──▼───▼───▼──┐
                    │ COLLECTOR   │
                    │ Vote / Merge│
                    │ / Synthesize│
                    └─────────────┘

When to use:     Independent perspectives on same problem
                 Latency-critical: wall clock = max(agent)
Cost profile:    N agents (parallel) + 1 aggregator
Failure mode:    API rate limits (N agents × requests > limit)
                 Race conditions: N agents have N(N-1)/2 potential conflicts
                 LLM-based synthesis can hallucinate false consensus
Mitigation:      Use explicit voting/weighted merge for deterministic aggregation
                 Cap at 5 parallel agents per dispatcher
```

**Red team application:** Multi-source data collection — SpiderFoot + Shodan + cert.sh + Amass simultaneously.

### 3.4 Handoff (Level 3)

```
┌─────────┐     ┌─────────┐     ┌─────────┐
│ Triage  │ ──► │ Recon   │ ──► │ Exploit │
│ Agent   │     │ Agent   │     │ Agent   │
│ (class.)│     │ (enumer)│     │ (breach)│
└─────────┘     └─────────┘     └─────────┘
     │                                               
     │ Only one active at a time                     
     │ Full control transfers with context           

When to use:     Optimal specialist emerges during processing
                 Cheapest multi-agent pattern
Cost profile:    1 active agent × chain depth
Failure mode:    INFINITE HANDOFF LOOPS — #1 production failure
                 Context loss compounds with each transfer
Mitigation:      Max handoff depth (3-5). Explicit handoff schema (JSON).
                 Transfer only structured context, not full conversation.
```

**Red team application:** Triage → intelligence gathering → exploit chain. Each stage hands off to next specialist.

### 3.5 Group Chat / Debate (Level 3)

```
┌─────────────────────────────────────────────────┐
│           SHARED CONVERSATION THREAD             │
│                                                  │
│ Agent A: "I found X. Recommend approach Y."     │
│ Agent B: "X is a false positive. Check Z."      │
│ Agent C: "Confirmed. Z has a known CVE."        │
│ Agent A: "Agreed. Path is Z via CVE-2026-XXXX." │
│                                                  │
│ Chat Manager controls: who speaks, when to stop  │
└─────────────────────────────────────────────────┘

When to use:     Quality verification, maker-checker loops
                 Requires multi-perspective validation
Cost profile:    N agents × rounds
                 Cap at 3 agents — beyond that, sycophancy cascades
Failure mode:    Conversation loops (agents never converge)
                 Sycophancy: agents agree with majority even when wrong
Mitigation:      Set max rounds (3-5). Maker uses cheap model,
                 checker uses capable model. Cost savings: 40-60%.
```

**Red team application:** Cross-validation of recon findings. One agent generates attack path, another validates, third checks for detection gaps.

### 3.6 Magentic / Adaptive Planning (Level 4)

```
┌─────────────────────────────────────────────────────┐
│ MANAGER AGENT                                        │
│ Builds TASK LEDGER (live, revised)                   │
│                                                      │
│ [ ] Subgoal 1: Recon network                                     │
│ [✓] Subgoal 2: Identify entry points                            │
│ [ ] Subgoal 3: Craft exploit payload (reassigned to Agent X)    │
│ [ ] Subgoal 4: Deploy persistence                               │
│ [ ] Subgoal 5: Exfil target data                                │
│                                                      │
│ Manager reorders, reassigns, adds as context evolves │
└─────────────────────────────────────────────────────┘

When to use:     Open-ended problems, no known solution path
                 Must discover approach during execution
Cost profile:    Unbounded until plan converges
Failure mode:    Slow to converge. Stalls on ambiguous goals.
                 Hard to estimate cost/time upfront.
Mitigation:      Set max iteration budget (token/cost cap).
                 After Nth iteration, escalate to human.
```

**Red team application:** Complex multi-stage campaigns where path depends on intermediate discoveries.

### 3.7 Blackboard (Level 4)

```
┌─────────────────────────────────────────────┐
│              BLACKBOARD (shared state)       │
│                                              │
│  key: "target.internal.ip"                   │
│  val: "10.0.1.55"                           │
│  key: "vulnerabilities"                     │
│  val: ["CVE-2026-XXXX on port 443"]         │
│  key: "credentials"                         │
│  val: {"user": "admin", "hash": "..."}      │
│  key: "task_status"                         │
│  val: {"recon": "done", "exploit": "fail" } │
└─────────────────────────────────────────────┘
      ▲           ▲           ▲           ▲
      │           │           │           │
┌─────┴──┐ ┌─────┴──┐ ┌─────┴──┐ ┌─────┴──┐
│ Recon  │ │Exploit │ │Pivot   │ │Exfil   │
│ Agent  │ │Agent   │ │Agent   │ │Agent   │
└────────┘ └────────┘ └────────┘ └────────┘

When to use:     Dynamic, event-driven collaboration
                 Next step depends on intermediate findings
Failure mode:    Write conflicts (two agents update same key)
                 Infinite loops (A writes → B triggers → A triggers)
Mitigation:      Optimistic locking (Redis). Max iteration counter per item.
                 Task-completion flags prevent re-processing.
```

**Red team application:** Multi-phase op where each agent writes findings to shared board and other agents react. Recon writes a discovered port → Exploit agent reads it and attacks → writes result → Pivot agent reads and moves laterally.

### 3.8 Swarm (Level 5)

```
┌──────────────────────────────────────────────────┐
│  No central orchestrator. Agents route to peers. │
│                                                   │
│         ┌─────────┐                              │
│         │ Agent A │◄────► Agent B                │
│         └────┬────┘      └────┬────┘             │
│              │                │                   │
│         ┌────▼────┐     ┌────▼────┐              │
│         │ Agent C │◄───►│ Agent D │              │
│         └─────────┘     └─────────┘              │
│                                                   │
│  Each agent knows its own capability set          │
│  Each agent routes to peer when out of depth      │
└──────────────────────────────────────────────────┘

When to use:     Highly unpredictable tasks, emergent problem-solving
Failure mode:    Near-impossible to debug. Non-deterministic routing.
                 Hard to predict cost or path.
Mitigation:      Not for production systems without extreme observability.
                 Use only for research/exploration.
```

**Red team application:** Experimental — autonomous hunt teams where recon agents discover surfaces and route to exploit agents dynamically.

---

## 4. Security Architecture

### 4.1 The Agentic Security Stack

```
┌─────────────────────────────────────────────────────────────┐
│ LAYER 5: GOVERNANCE & POLICY                                │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ RBAC · ABAC · OPA policies · Data classification        │ │
│ │ Per-workflow token budgets · Cost limits                 │ │
│ │ Compliance evidence (SOC 2, HIPAA, PCI)                 │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 4: ORCHESTRATION SECURITY                             │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Supervisor agent hardening · Planner constraints        │ │
│ │ Delegation depth limits · Max iteration budgets         │ │
│ │ Human-in-the-loop gates for destructive actions         │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 3: AGENT IDENTITY & AUTHENTICATION                    │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ AIP (Agent Identity Protocol) — cryptographic identity │ │
│ │ Delegation tokens (Biscuit/macaroon) — attenuate scope │ │
│ │ OAuth 2.1 + PKCE — every agent has a verifiable ID     │ │
│ │ Token: actor + subject claims, scope narrows per hop   │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 2: MCP SECURITY                                       │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ MCP Gateway — central authN/Z, policy, rate limits      │ │
│ │ STDIO hardening — minimal container, non-root, R/O FS   │ │
│ │ Supply chain — version pinning, Sigstore verification   │ │
│ │ Tool poisoning detection — audit tool descriptions      │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 1: EXECUTION SANDBOX                                  │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Firecracker / gVisor micro-VMs for agent execution      │ │
│ │ seccomp / SELinux profiles per agent type               │ │
│ │ Network egress allowlists — deny all by default         │ │
│ │ Read-only filesystem for all non-writing agents         │ │
│ └─────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
```

### 4.2 The "Config is Code" Principle (from Paperclip CVE)

The foundational security insight from CVE-2026-41679: **agent configuration IS executable code.**

```
Every one of these must be treated as code:
  ├── YAML/JSON agent definitions
  ├── MCP server STDIO command strings
  ├── Tool descriptions (name, schema, instructions)
  ├── Prompt templates
  ├── Parameter schemas
  ├── Agent skill files (SKILL.md)
  ├── Orchestration workflow definitions
  ├── Inter-agent handoff schemas
  └── Role/backstory prompts

Security posture:
  - Code review for every config change
  - Version control (git) for all definitions
  - Sign config artifacts (Sigstore/Cosign)
  - Pin versions — never deploy "latest"
  - Audit trail on every config modification
```

### 4.3 MCP Security Hardening (Expanded)

From the Coalition for Secure AI RSAC 2026 findings, Knostic scan of ~2,000 MCP servers (all lacked auth), and the Microsoft MCP Security 2026 report:

```
┌────────────────────┬─────────────────────────────────────────┐
│ Risk               │ Mitigation                               │
├────────────────────┼─────────────────────────────────────────┤
│ Prompt injection   │ Validate all tool descriptions + params  │
│ (tool poisoning)   │ Separate trusted/untrusted MCP servers   │
│                    │ Never craft prompts from user content    │
├────────────────────┼─────────────────────────────────────────┤
│ No auth on MCP     │ OAuth 2.1 + PKCE per server              │
│ servers            │ No long-lived tokens. Rotate every hour. │
│ (Knostic finding:  │ Token audience bound to specific server  │
│ ~2,000 servers,    │                                         │
│ 0 with auth)       │                                         │
├────────────────────┼─────────────────────────────────────────┤
│ Supply chain rug   │ Pin versions. Verify build provenance.   │
│ pull               │ Watch for ownership changes on upstream  │
│                    │ SBOM per MCP server                      │
├────────────────────┼─────────────────────────────────────────┤
│ Shadow MCP         │ Inventory all MCP servers. Unregistered  │
│ servers            │ = blocked by default at gateway.         │
│                    │ Auto-discover via network scan.          │
├────────────────────┼─────────────────────────────────────────┤
│ Over-permissioned  │ Narrow OAuth scopes. "Delete" scope only │
│ tokens             │ when agent needs to delete. Audited.     │
│                    │ Scoped per resource, not per server.     │
├────────────────────┼─────────────────────────────────────────┤
│ Multi-hop identity │ AIP (Agent Identity Protocol).           │
│ chain loss         │ Delegation tokens with attenuated scope. │
│                    │ Policy evaluated at every hop.           │
└────────────────────┴─────────────────────────────────────────┘
```

### 4.4 The Blast Radius Principle

```
┌───────────────────────────────────────────────────────────┐
│ BLAST RADIUS = every agent can reach                     │
│                  ├── These MCP servers                    │
│                  ├── These tools                           │
│                  ├── This memory/state                     │
│                  ├── These peer agents                     │
│                  └── These filesystem paths               │
│                                                           │
│ RULE: Narrow the blast radius of every agent to the       │
│       absolute minimum required for its role.             │
│                                                           │
│ Example: Recon agent needs:                               │
│   ✓ web_search() tool                                     │
│   ✗ write_file() tool                                     │
│   ✗ db_query() tool                                       │
│   ✗ access to peer agents                                 │
│   ✗ access to credential store                            │
│                                                           │
│ Exploit agent needs:                                      │
│   ✓ command execution (sandboxed)                         │
│   ✓ write_file() to output directory                     │
│   ✓ A2A comms with Pivot agent                            │
│   ✗ access to production database                         │
│   ✗ ability to exfiltrate to arbitrary endpoints          │
└───────────────────────────────────────────────────────────┘
```

---

## 5. State & Memory Architecture

### 5.1 Three-Tier Memory Model

```
┌────────────────────────────────────────────────────────────┐
│ TIER 1: WORKING MEMORY (ephemeral, per-task)              │
│   Storage: In-context (LLM context window)                │
│   Lifetime: Single agent call                             │
│   Contents: Current reasoning, tool outputs, partial state │
│   Risk: Context window overflow                            │
│   Mitigation: Token budgets, message compression           │
├────────────────────────────────────────────────────────────┤
│ TIER 2: SESSION MEMORY (persistent, per-conversation)     │
│   Storage: State store (Redis, SQLite, file)              │
│   Lifetime: Entire agent session (hours to days)          │
│   Contents: Conversation history, completed tasks,         │
│             intermediate findings, credentials cache       │
│   Key property: Survives agent restarts, context pruning   │
│   Mitigation: Session expiry, max size limits, PII redact  │
├────────────────────────────────────────────────────────────┤
│ TIER 3: SEMANTIC MEMORY (long-term, cross-session)        │
│   Storage: Vector DB / knowledge graph (Chroma, pgvector) │
│   Lifetime: Months to permanent                           │
│   Contents: Learned facts, reusable procedures,            │
│             skill definitions, absorbed knowledge          │
│   Key property: Injected into new sessions as context      │
│   Mitigation: Trust metadata, provenance tracking,         │
│               periodic consolidation, corruption detection │
└────────────────────────────────────────────────────────────┘
```

### 5.2 Orchestration State Management

```
┌──────────────────────────────────────────────────────────┐
│ ORCHESTRATOR STATE (per workflow execution)              │
│                                                          │
│  {                                                        │
│    "workflow_id": "recon-20261003-001",                  │
│    "status": "in_progress",                               │
│    "plan": {                                              │
│      "phases": ["recon", "exploit", "exfil"],            │
│      "current_phase": "recon",                           │
│      "agents_assigned": {                                 │
│        "recon": ["spiderfoot_mcp", "amass_agent"],       │
│        "exploit": null,                                   │
│        "exfil": null                                      │
│      }                                                    │
│    },                                                     │
│    "artifacts": {                                         │
│      "domains_found": ["target.com", "admin.target.com"],│
│      "ports_open": [80, 443, 8443],                      │
│      "cves_possible": ["CVE-2026-XXXX"]                  │
│    },                                                     │
│    "errors": [],                                          │
│    "checkpoints": [                                       │
│      {"phase": "recon", "time": "2026-10-03T02:00:00Z",  │
│       "state_hash": "abc123"}                            │
│    ]                                                      │
│  }                                                        │
│                                                          │
│ Key properties:                                           │
│   - Checkpointed at every phase (survives failure)       │
│   - State hash for integrity verification                │
│   - Error log with rollback capability                   │
│   - Artifacts accumulate across phases                   │
└──────────────────────────────────────────────────────────┘
```

### 5.3 Inter-Agent Context Handoff

When agents hand off tasks, pass **structured schema, not conversation history:**

```
BAD: Full conversation dump (context expensive, LLM distills lossily)
  Agent A → Agent B: [50K tokens of raw chat history]

GOOD: Structured handoff object (schema-defined, compressed, lossless)
  Agent A → Agent B: {
    "task": "exploit_port_443",
    "findings": {
      "target_ip": "10.0.1.55",
      "port": 443,
      "service": "nginx 1.24.0",
      "cve": "CVE-2026-XXXX",
      "confidence": 0.85,
      "source": "nmap-agent"
    },
    "context_chain": ["recon", "scan", "identify"],
    "constraints": {
      "max_duration_s": 300,
      "destructive_allowed": false
    }
  }
```

---

## 6. Framework Selection Guide

### 6.1 Comparison Matrix

```
┌──────────────────┬───────────┬───────────┬────────────┬───────────┬──────────────┐
│ Framework        │ Best For  │ Orchestr. │ State Mgmt │ Security  │ MCP Support  │
│                  │           │ Model     │            │ First?    │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ LangGraph        │ Complex   │ DAG /     │ Checkpoint │ No (add   │ Full (native)│
│                  │ stateful  │ State     │ per node,  │ OPA/      │              │
│                  │ workflows │ Graph     │ time-travel│ gateway)  │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ CrewAI           │ Role-     │ Crews /   │ Task-based │ No (add   │ Full         │
│                  │ based     │ Hierarch. │ sequential │ gateway)  │              │
│                  │ business  │           │ memory     │           │              │
│                  │ processes │           │            │           │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ MS Agent         │ Enterprise│ Graph /   │ Session    │ Yes       │ Full (native)│
│ Framework        │ Azure     │ GroupChat │ state +    │ (RBAC,    │              │
│                  │ stack     │ Handoff   │ checkpoint │ OAuth)    │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ Google ADK      │ Google    │ Hierarch. │ Session    │ Partial   │ Full + A2A   │
│                  │ Cloud     │ tree +    │ (per agent)│           │ (donated to  │
│                  │ env       │ A2A       │            │           │ LF)          │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ OpenAI Agents    │ Tuned to  │ Handoff   │ External   │ No (add   │ Full         │
│ SDK              │ OpenAI    │ (tool     │ (BYO DB)  │ gateway)  │              │
│                  │ models    │ call)     │            │           │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ PyAgent          │ 18        │ All 4     │ 3-tier     │ Yes (OPA)│ Full         │
│                  │ patterns  │ tiers     │ memory     │           │              │
│                  │ out-of-box│           │            │           │              │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ Hermes Agent     │ CLI agent │ None      │ Memory /   │ Limited   │ Full         │
│ (this runtime)   │ lifecycle │ (handoff  │ Session    │ (profiles)│              │
│                  │           │ via       │ search     │           │              │
│                  │           │ delegate) │            │           │              │
└──────────────────┴───────────┴───────────┴────────────┴───────────┴──────────────┘
```

### 6.2 Selection Decision Tree

```
┌──────────────────────────────────────────────────────────────┐
│ Are you on a specific cloud stack?                           │
│   Microsoft / Azure   → MS Agent Framework                    │
│   Google Cloud        → Google ADK                            │
│   AWS / self-hosted   → LangGraph or PyAgent                  │
│   None / multi-cloud  → LangGraph or CrewAI                   │
├──────────────────────────────────────────────────────────────┤
│ Is security-first governance required?                       │
│   YES → MS Agent Framework (Entra ID) or PyAgent (OPA built-in)│
│         Add MCP Gateway layer regardless of framework choice  │
├──────────────────────────────────────────────────────────────┤
│ Quick prototype needed?                                       │
│   Role-based → CrewAI (fastest to working multi-agent)       │
│   Tool-heavy → OpenAI Agents SDK (minimal abstraction)        │
├──────────────────────────────────────────────────────────────┤
│ Production reliability critical?                              │
│   Deterministic DAG → LangGraph (checkpointing, time-travel) │
│   Complex routing  → Conductor (Microsoft, Jinja2 routing)   │
├──────────────────────────────────────────────────────────────┘
```

---

## 7. Red Team Integration Mapping

### 7.1 Pattern-to-Kill-Chain Mapping

```
┌────────────────────────────┬─────────────────────────────────┐
│ Red Team Phase             │ Optimal Orchestration Pattern   │
├────────────────────────────┼─────────────────────────────────┤
│ Reconnaissance (TA0043)    │ Concurrent / Fan-Out            │
│                            │ SpiderFoot + Amass + Shodan     │
│                            │ in parallel → collector merges  │
├────────────────────────────┼─────────────────────────────────┤
│ Resource Dev (TA0042)      │ Sequential pipeline             │
│                            │ Domain buy → DNS setup →        │
│                            │ cert → VPS provision             │
├────────────────────────────┼─────────────────────────────────┤
│ Initial Access (TA0001)    │ Handoff                         │
│                            │ Triage → phish → exploit chain  │
├────────────────────────────┼─────────────────────────────────┤
│ Execution (TA0002)         │ Supervisor-worker               │
│                            │ Deploy dropper, stager, payload │
├────────────────────────────┼─────────────────────────────────┤
│ Persistence (TA0003)       │ Blackboard                      │
│                            │ Agents monitor and react to     │
│                            │ access state changes            │
├────────────────────────────┼─────────────────────────────────┤
│ PrivEsc (TA0004)           │ Handoff                         │
│                            │ Enum → attempt → next technique │
├────────────────────────────┼─────────────────────────────────┤
│ Credential Access (TA0006) │ Concurrent                      │
│                            │ LSASS + browser + Kerberoast    │
│                            │ all parallel → merge findings   │
├────────────────────────────┼─────────────────────────────────┤
│ Lateral Movement (TA0008)  │ Blackboard                      │
│                            │ Each agent writes pivot points  │
│                            │ Others read and exploit         │
├────────────────────────────┼─────────────────────────────────┤
│ Exfiltration (TA0010)      │ Sequential pipeline             │
│                            │ Collect → archive → encrypt →   │
│                            │ exfil → verify                  │
├────────────────────────────┼─────────────────────────────────┤
│ Reporting                  │ Supervisor-worker               │
│                            │ Planner assigns sections to     │
│                            │ specialist writer agents         │
└────────────────────────────┴─────────────────────────────────┘
```

### 7.2 Agent Roles in a Red Team MAS

```
┌────────────────┬────────────────────────────┬─────────────────┐
│ Agent Role     │ Responsibilities            │ MCP Tools       │
├────────────────┼────────────────────────────┼─────────────────┤
│ Recon Agent    │ OSINT, surface mapping,    │ web_search,     │
│                │ subdomain enumeration,     │ web_extract,    │
│                │ technology profiling       │ SpiderFoot API  │
├────────────────┼────────────────────────────┼─────────────────┤
│ Intrusion Agent│ Phishing, exploit dev,     │ SMTP relay,     │
│                │ payload generation,        │ payload builder,│
│                │ initial access             │ browser         │
├────────────────┼────────────────────────────┼─────────────────┤
│ C2 Agent       │ Beacon deployment,         │ Sliver/Mythic   │
│                │ heartbeat management,      │ API, MCP -> C2  │
│                │ command dispatch           │ bridge          │
├────────────────┼────────────────────────────┼─────────────────┤
│ Pivot Agent    │ Lateral movement,          │ impacket, SSH,  │
│                │ AD exploitation, cloud      │ BloodHound API  │
│                │ escalation                 │                 │
├────────────────┼────────────────────────────┼─────────────────┤
│ Intel Agent    │ Credential parsing,        │ Mimikatz API,   │
│                │ data collection,            │ hash parser,    │
│                │ intelligence aggregation   │ data archiver   │
├────────────────┼────────────────────────────┼─────────────────┤
│ Report Agent   │ Findings aggregation,      │ Markdown writer,│
│                │ ATT&CK mapping,            │ ATT&CK Nav API, │
│                │ report generation          │ template engine │
├────────────────┼────────────────────────────┼─────────────────┤
│ Watchdog Agent │ OPSEC monitoring,          │ log parser,     │
│                │ detection checks,          │ network scanner,│
│                │ kill switch trigger        │ alert handler   │
└────────────────┴────────────────────────────┴─────────────────┘
```

### 7.3 Novel: MCP Bridge for C2 Operations

Building on SANS SEC565 (2026) research — using MCP servers as a C2 transport layer:

```
┌─────────────────────────────────────────────────────────────┐
│ MCP Bridge C2 Architecture                                  │
│                                                             │
│  ATTACKER SIDE                    TARGET SIDE               │
│  ┌────────────────┐              ┌──────────────────┐      │
│  │ C2 Orchestrator│              │ MCP Agent Runtime│      │
│  │ Agent          │              │                  │      │
│  │                │  MCP STDIO  │  "tools/call"     │      │
│  │ execute cmd ───┼─────────────►  execute(name:     │      │
│  │                │              │    "shell_cmd")   │      │
│  │                │◄─────────────┤  return: output   │      │
│  │ read output    │              │                  │      │
│  └────────────────┘              └──────────────────┘      │
│                                                             │
│  The MCP STDIO transport is indistinguishable from         │
│  legitimate agent tool usage to network monitoring.         │
│  C2 commands are JSON-RPC "tools/call" messages.           │
│                                                             │
│  RISK: This is novel (2026). Detection coverage is low.    │
│  ADVANTAGE: Blends with normal agent traffic on the host.  │
│  DEFENSE: Monitor for suspicious tool definitions in MCP   │
│           server configs. Audit all STDIO server commands.  │
└─────────────────────────────────────────────────────────────┘
```

---

## 8. Observability & Audit

### 8.1 What to Trace

Every agent interaction must produce structured logs:

```
┌─────────────────────────────────────────────────────────────┐
│ AGENT TRACE RECORD                                          │
│ {                                                            │
│   "trace_id": "recon-20261003-001",                         │
│   "span_id": "spiderfoot-mcp-call-003",                     │
│   "parent_span": "supervisor-plan-001",                      │
│   "agent": "recon-agent-v1",                                │
│   "action": "tools/call",                                   │
│   "tool": "web_search",                                     │
│   "input_hash": "sha256:abc...",                            │
│   "output_hash": "sha256:def...",                           │
│   "tokens_in": 452,                                          │
│   "tokens_out": 128,                                         │
│   "cost_usd": 0.0042,                                        │
│   "duration_ms": 3402,                                       │
│   "status": "success",                                       │
│   "timestamp": "2026-10-03T02:00:00Z",                      │
│   "model": "claude-opus-4.6-1m"                             │
│ }                                                            │
│                                                              │
│ Collect into: OpenTelemetry-compatible trace store            │
│ Visualize in: LangSmith, Weights & Biases, custom dashboard  │
│ Query via: trace_id across spans → full execution timeline   │
└─────────────────────────────────────────────────────────────┘
```

### 8.2 Cost Tracking Per Pattern

```
Set per-workflow budgets BEFORE execution:

  Sequential (4 agents):    Max 50K tokens / $0.15
  Supervisor (1+3):         Max 80K tokens / $0.25
  Concurrent (4 agents):    Max 60K tokens / $0.18
  Handoff (3 hops):         Max 40K tokens / $0.12
  Group chat (3×3 rounds):  Max 90K tokens / $0.28
  Magentic:                 CIRCUIT BREAKER at 200K tokens

Gateways enforce budgets. Exceeded = circuit breaker + notification.
```

### 8.3 Audit Trail Requirements

```
Non-negotiable for red team agent ops:

  [✓] Every tool call logged with input_hash and output_hash
  [✓] Every agent handoff logged with full context summary
  [✓] Every MCP server connection logged (source agent, server, transport)
  [✓] Every configuration change logged (who, what, when, diff)
  [✓] Every destructive action requires HITL gate + audit record
  [✓] Token/cost budget per workflow, per agent, per session
  [✓] Trace ID chains: user request → orchestrator → agent → MCP call
```

---

## 9. Deployment Patterns

### 9.1 Deterministic vs Dynamic Orchestration

```
┌──────────────────────┬──────────────────────────┬────────────────┐
│                      │ DETERMINISTIC (DAG)      │ DYNAMIC (LLM   │
│                      │                         │ routing)       │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Topology             │ Fixed at design time     │ Emerges at     │
│                      │                          │ runtime        │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Debugging            │ Trivial — step N fails   │ Hard — non-    │
│                      │ = agent N has the bug    │ deterministic  │
│                      │                          │ execution path │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Cost                 │ Predictable (N × tokens) │ Variable (may  │
│                      │                          │ loop/explode) │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Best for             │ Known workflows, batch   │ Discovery,     │
│                      │ processing, ETL          │ research, open │
│                      │                          │ problems       │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Blast radius         │ Bounded by design        │ Uncapped       │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Audit ability        │ Full traceability        │ Partial (LLM   │
│                      │                          │ decision not   │
│                      │                          │ deterministic) │
└──────────────────────┴──────────────────────────┴────────────────┘

RECOMMENDATION: Default to DAG. Reserve dynamic for the specific
sub-problem that genuinely requires it.
```

### 9.2 Hybrid Patterns (Production Norm)

Production systems rarely use one pure pattern. Common hybrid:

```
┌──────────────────────────────────────────────────────────┐
│ SUPERVISOR at top level                                  │
│  │                                                       │
│  ├── PIPELINE for stage 1 (recon: fixed sequence)        │
│  │    SpiderFoot → Amass → Nmap → Vuln scan              │
│  │                                                       │
│  ├── CONCURRENT for stage 2 (multi-source collection)    │
│  │    cert.sh + Shodan + GitHub dorking in parallel      │
│  │                                                       │
│  ├── HANDOFF for stage 3 (triage → exploit chain)        │
│  │    Triage agent → Exploit agent → Pivot agent         │
│  │                                                       │
│  └── SUPERVISOR for stage 4 (report synthesis)           │
│       Writers generate sections in parallel              │
│       Supervisor assembles final report                  │
│                                                          │
│ Each stage gets its own pattern.                         │
│ The supervisor decides which stage to run next.          │
└──────────────────────────────────────────────────────────┘
```

---

## 10. Hardening Checklist

### 10.1 Pre-Deployment Security Gates

```
[ ] Every agent definition code-reviewed
[ ] Every MCP server version-pinned + provenance verified
[ ] Every STDIO server command string audited (is it executable?)
[ ] OAuth 2.1 + PKCE configured on every MCP server
[ ] Tokens scoped per-resource (never wildcard)
[ ] Token rotation policy set (max 1 hour for operational agents)
[ ] Gateway in path — no agent talks directly to MCP server
[ ] Network egress deny-by-default for sandboxed agents
[ ] Read-only filesystem for all non-writing agents
[ ] Per-workflow token/cost budgets defined
[ ] Circuit breakers configured (max iterations, max cost)
[ ] Human-in-the-loop gates for destructive actions
[ ] All agent interactions traced to OpenTelemetry store
[ ] Trace IDs chainable from user request to individual tool call
[ ] Shadow MCP detection active (inventory scan)
[ ] Supply chain monitoring on all MCP server dependencies
[ ] Agent identity attested (AIP or similar)
[ ] Delegation scope narrows at every hop (never expands)
```

### 10.2 Runtime Monitoring

```
┌──────────────────────────┬──────────────────────────────┐
│ What to Monitor          │ Alert Threshold              │
├──────────────────────────┼──────────────────────────────┤
│ Orchestration loops      │ >5 iterations without progress│
│ Agent handoff depth      │ >5 hops without completion   │
│ Token consumption/workfl.│ >2x estimated budget         │
│ Cost/workflow            │ >$0.50 USD                   │
│ MCP server response      │ >30s timeout                 │
│ Agent context window     │ >90% of limit                │
│ Inter-agent latency      │ >10s between handoffs        │
│ Tool call failure rate   │ >10% of calls                │
│ Unregistered MCP servers │ Any = immediate              │
│ Long-lived tokens        │ >1 hour without rotation     │
│ Unattenuated delegation  │ Scope non-narrowing at hop   │
└──────────────────────────┴──────────────────────────────┘
```

---

## 11. Tool & Resource Reference

### 11.1 Frameworks (2026)

| Framework | URL | License | Key Strength |
|-----------|-----|---------|--------------|
| LangGraph | github.com/langchain-ai/langgraph | MIT | Stateful DAG, checkpointing |
| CrewAI | github.com/crewAIInc/crewAI | MIT | Fast role-based prototype |
| MS Agent Framework | github.com/microsoft/agent-framework | MIT | Enterprise, Azure, security |
| Google ADK | github.com/google/adk | Apache 2.0 | A2A protocol, GCP native |
| OpenAI Agents SDK | github.com/openai/openai-agents-python | MIT | Clean handoff, MCP native |
| PyAgent | github.com/pyagent-core/pyagent | MIT | 18 patterns, OPA built-in |
| Conductor | opensource.microsoft.com | MIT | Deterministic, Jinja2 routing |

### 11.2 MCP Ecosystem

| Tool | Purpose | Security Note |
|------|---------|---------------|
| MCP Gateway | Centralized authN/Z, policy, audit | REQUIRED for production |
| ScaleMCP | Dynamic tool inventory sync | Needs gateway integration |
| AgentMaster | MCP + A2A integration comoposer | Experimental |
| MCP Inspector | Debug/audit MCP server configs | Scan for STDIO injection |
| MCP OWASP Top 10 | Risk framework for MCP | Maps to Azure/OPA controls |

### 11.3 Observability

| Tool | Purpose |
|------|---------|
| LangSmith | Framework-agnostic tracing, eval |
| OpenTelemetry | Standard collector and export |
| Weights & Biases | Experiment tracking, cost analysis |
| VECTR | Red team engagement tracking |
| ATT&CK Navigator | Technique coverage visualization |

---

## 12. The Paperclip Lesson (Reinforced)

CVE-2026-41679 taught us: **"Configuration is code."** This principle applies at every level of agent orchestration:

```
    Agent YAML/JSON config       → IS CODE
    MCP server STDIO command     → IS CODE
    Tool description schema      → IS CODE (tool poisoning)
    Inter-agent handoff schema   → IS CODE
    Orchestration workflow def   → IS CODE
    Skill file (SKILL.md)        → IS CODE
    Prompt template              → IS CODE
    Role/backstory assignment    → IS CODE (prompt injection)

    Every one of these can be exploited.
    Every one of these must be reviewed, signed, pinned, and audited.
```

**The attack chain that works against any orchestrated system:**

```
1. Attacker submits crafted input to an agent
2. Agent treats input as data — but orchestrator treats it as workflow
3. Input becomes plan → determines which agent runs, with what tools
4. Orchestrator delegates to a privileged agent
5. Privileged agent executes attacker's intended action

⚠ This is not theoretical. This is Paperclip applied to orchestration.
⚠ Mitigation: Validate at every hop. Never trust inputs as workflow.
```

---

**End of ptSlick Agentic Orchestrated Blueprint v1.0**

> "Start with a single agent. Orchestrate only when the limitation forces it.
>  Monitor every handoff. Audit every config. Budget every token."