📖 Blueprints & Methodologies
Pentest methodology documents, operational blueprints, and security references.
← Back to list
agentic-orchestrated-blueprint.md
agentic-orchestrated-blueprint.md
# ptSlick — Agentic Orchestrated Blueprint v1.0
**Classification**: Internal — SlickLab Agent Operations
**Build**: 2026-10-03 | **Next Review**: 2026-11-03
**Author**: Colnex (SlickLab Operational Intelligence)
---
## 0. Executive Summary
This blueprint defines how to **orchestrate multiple AI agents into coordinated, auditable, security-hardened workflows** for red team operations and security research. It covers 8 orchestration patterns, 2 communication protocols, 3 memory tiers, and the security architecture that binds them together — informed by the Paperclip CVE, MCP STDIO injection, and the 2026 agent security landscape.
**Key insight from research (2026 consensus):** Start with a single agent. Only add multi-agent orchestration when you hit a specific limitation — tool overload (15-20+ tools), context window overflow, distinct security boundaries, or genuine need for specialization. Multi-agent multiplies token cost roughly **4-15x** over single-agent. The patterns in this blueprint each pay that tax differently.
---
## 1. Architecture Principles
### 1.1 The Hierarchy of Orchestration Complexity
```
Cost ←───────────────────────────────────────→ Complexity
SINGLE AGENT LEVEL 0
└─ One agent, one context, N tools
SEQUENTIAL PIPELINE LEVEL 1
└─ Agent A → Agent B → Agent C
Deterministic, fixed order, every step depends on previous
SUPERVISOR / ORCHESTRATOR-WORKER LEVEL 2
└─ Lead agent decomposes → delegates → synthesizes
Dynamic subtasks, specialist workers, single accountability
CONCURRENT / FAN-OUT/FAN-IN LEVEL 2
└─ Same input → N parallel agents → aggregated
Time win: wall clock = max(agents), not sum
HANDOFF LEVEL 3
└─ Task passes agent-to-agent, one active at a time
Cheapest multi-agent pattern (no parallel context)
GROUP CHAT / DEBATE LEVEL 3
└─ Shared thread, chat manager controls turn order
Best for maker-checker quality loops, capped at 3
MAGENTIC / ADAPTIVE PLANNING LEVEL 4
└─ Manager builds live task ledger, iterates
For problems with no predetermined solution path
BLACKBOARD LEVEL 4
└─ Shared state store, agents react to events
Event-driven, needs conflict resolution
SWARM LEVEL 5
└─ Peer-to-peer routing, no central coordinator
Highest complexity, hardest to debug
```
**Rule:** Default to LEVEL 0. Only escalate when you can articulate the specific limitation that forces it.
### 1.2 When to Orchestrate (and When Not To)
```
QUESTION CHAIN:
┌──────────────────────────────────────────────────────┐
│ Can one agent handle this in one context window? │
│ YES → Single agent. Done. │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Is the task decomposable into fixed stages? │
│ YES → Sequential pipeline (Level 1) │
│ Cost = N agents × 1 pass │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Are the subtasks known but independent? │
│ YES → Supervisor/orchestrator-worker (Level 2) │
│ Cost = 1 planner + N workers + 1 synthesizer │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Do you need multiple perspectives on one input? │
│ YES → Concurrent/fan-out (Level 2) │
│ Cost = N agents + 1 aggregator │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Is the optimal routing unknown until runtime? │
│ YES → Handoff (Level 3) │
│ Cost = 1 active agent × depth │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Do you need quality verification / reflection? │
│ YES → Group chat / debate (Level 3) │
│ Cost = N agents × rounds, cap at 3 agents │
│ NO ↓ │
├──────────────────────────────────────────────────────┤
│ Is the solution path unknowable upfront? │
│ YES → Magentic / adaptive planning (Level 4) │
│ Cost = unbounded until plan converges │
│ NO → You probably don't need multi-agent │
└──────────────────────────────────────────────────────┘
```
### 1.3 Critical Cost Data (2026 Benchmarks)
```
┌────────────────────────────────┬──────────┬─────────────┐
│ Pattern │ Token │ Coordination│
│ │ Multiple │ Overhead per │
│ │ │ step │
├────────────────────────────────┼──────────┼─────────────┤
│ Single agent with tools │ ~4x │ 0ms │
│ Sequential pipeline (4 agents) │ ~29k │ ~950ms │
│ Supervisor-worker (5 workers) │ ~15x │ ~2-3s │
│ Concurrent (4 agents) │ ~4x │ ~500ms │
│ Group chat (3 agents, 3 rnds) │ ~9x │ ~4-8s │
│ Magentic (unbounded) │ ~20x+ │ variable │
│ Swarm (N agents × rounds) │ N×R×2 │ N(N-1)/2 │
└────────────────────────────────┴──────────┴─────────────┘
Source: Microsoft Azure Architecture Center, Anthropic engineering data, 2026.
Single-agent-with-tools baseline: ~4x standard chat. Multi-agent starts at ~15x.
```
---
## 2. Communication Protocols
### 2.1 Protocol Topology
```
┌─────────────────────┐
│ USER/OPERATOR │
└──────────┬──────────┘
│
┌──────────▼──────────┐
│ ORCHESTRATOR │
│ (Agent Manager) │
└──┬────────────┬─────┘
│ │
┌────────▼──┐ ┌────▼────────┐
│ MCP CLIENT│ │ A2A CLIENT │
│ (tools) │ │ (peers) │
└───┬───────┘ └────┬────────┘
│ │
┌────────▼────────┐ ┌───▼──────────────┐
│ MCP Servers │ │ Peer Agents │
│ ┌────────────┐ │ │ ┌────────────┐ │
│ │ Search │ │ │ │ Code Agent │ │
│ │ DB Query │ │ │ │ Research │ │
│ │ File Sys │ │ │ │ Deploy │ │
│ │ API Gate │ │ │ │ Analyst │ │
│ │ Sandbox │ │ │ └────────────┘ │
│ └────────────┘ │ └──────────────────┘
└─────────────────┘
MCP = Model Context Protocol (tool access, client-server)
A2A = Agent-to-Agent protocol (peer delegation, HTTP/JSON-RPC)
```
### 2.2 Model Context Protocol (MCP) Architecture
MCP is the **USB interface for AI agents** — it standardizes how agents connect to tools, data sources, and services.
```
┌─────────────────────────────────────────────────────────┐
│ MCP Host (Orchestrator / Agent Runtime) │
│ ┌──────────────────────────────────────────────────┐ │
│ │ MCP Client: Agent ↔ Server communication │ │
│ │ - tools/call (execute tool by name) │ │
│ │ - resources/read (access data by URI) │ │
│ │ - prompts/get (retrieve prompt templates) │ │
│ │ - session management (stateful exchanges) │ │
│ └──────────────────────────────────────────────────┘ │
│ │ │ │
│ ┌────────▼────────┐ ┌──────▼───────────────┐ │
│ │ MCP Server A │ │ MCP Server B │ │
│ │ Search Tools │ │ Database Tools │ │
│ │ ─────────────── │ │ ───────────────────── │ │
│ │ web_search() │ │ query_db() │ │
│ │ web_extract() │ │ read_table() │ │
│ │ url_scan() │ │ write_row() │ │
│ └─────────────────┘ └───────────────────────┘ │
└─────────────────────────────────────────────────────────┘
```
**MCP Transport Types (critical security dimension):**
| Transport | How It Works | Security Risk |
|-----------|-------------|---------------|
| **STDIO** | Server runs as subprocess, communicates over stdin/stdout | **HIGH** — CVE-2026-30623, CVE-2026-30615. STDIO server definition is executable content. 200k+ instances affected, 150M+ package downloads. Tool descriptions parsed as code. |
| **HTTP/SSE** | Server listens on HTTP port, agent sends JSON-RPC | **MEDIUM** — Network exposure, needs auth. OAuth 2.1 adopted 2026. |
| **WebSocket** | Persistent bidirectional channel | **MEDIUM** — Same as HTTP + session management |
**MCP Security Hardening (from Paperclip + STDIO lessons):**
```
┌────────────────┬────────────────────────────────────────────────┐
│ Rule │ Implementation │
├────────────────┼────────────────────────────────────────────────┤
│ 1. Config is │ Every MCP server definition is executable. │
│ code │ Treat tool descriptions, parameter schemas, │
│ │ and STDIO command strings as code. Audit them. │
├────────────────┼────────────────────────────────────────────────┤
│ 2. Narrow │ One MCP server per capability. Never a "super │
│ servers │ server" with every tool. Scope = blast radius. │
├────────────────┼────────────────────────────────────────────────┤
│ 3. Least │ OAuth 2.1 + PKCE. No long-lived tokens. │
│ privilege │ Scopes narrow per resource. Token rotation. │
├────────────────┼────────────────────────────────────────────────┤
│ 4. Gateway │ Route all MCP traffic through gateway. Central │
│ pattern │ authN/Z, policy-as-code (OPA), rate limits, │
│ │ audit. Kill switch per server. │
├────────────────┼────────────────────────────────────────────────┤
│ 5. No STDIO │ Default to HTTP/SSE or WebSocket transport. │
│ by default │ STDIO only when server on same host, non-root, │
│ │ read-only filesystem, minimal container. │
├────────────────┼────────────────────────────────────────────────┤
│ 6. Supply │ Pin server versions. Verify build provenance │
│ chain │ (Sigstore). Watch for ownership changes. │
│ security │ Treat MCP servers like npm/pip dependencies. │
├────────────────┼────────────────────────────────────────────────┤
│ 7. Shadow MCP │ Register all MCP servers. Inventory them. │
│ detection │ Unregistered servers = ungoverned = risk. │
└────────────────┴────────────────────────────────────────────────┘
```
### 2.3 Agent-to-Agent Protocol (A2A)
A2A governs peer-level agent communication — delegation, negotiation, capability discovery.
```
┌────────────────────────────────────────────────────────────┐
│ A2A Communication Flow │
│ │
│ Agent A (Requester) Agent B (Responder) │
│ ┌──────────────────┐ ┌──────────────────┐ │
│ │ A2A Agent Card │ ───────► │ A2A Agent Card │ │
│ │ - capabilities │ GET │ - capabilities │ │
│ │ - auth methods │ │ - auth methods │ │
│ │ - endpoint URL │ │ - endpoint URL │ │
│ └───────┬──────────┘ └───────┬──────────┘ │
│ │ │ │
│ │ POST /task (JSON-RPC) │ │
│ ├────────────────────────────►│ │
│ │ { │ │
│ │ "jsonrpc": "2.0", │ │
│ │ "method": "tasks/send", │ │
│ │ "params": { │ │
│ │ "id": "task-001", │ │
│ │ "message": { │ │
│ │ "role": "user", │ │
│ │ "parts": [...] │ │
│ │ } │ │
│ │ } │ │
│ │ } │ │
│ │◄────────────────────────────┤ │
│ │ { "id": "task-001", │ │
│ │ "status": "accepted" } │ │
│ │ │ │
│ │ (Agent B processes...) │ │
│ │ │ │
│ │◄────────────────────────────┤ │
│ │ POST /task (update) │ │
│ │ { "status": "completed", │ │
│ │ "artifact": {...} } │ │
│ └─────────────────────────────┘ │
└────────────────────────────────────────────────────────────┘
Security properties:
- Cryptographic signing at every hop
- Delegation scope MUST narrow (never expand)
- Identity attestation via AIP (Agent Identity Protocol)
- Every A2A exchange logged to orchestration state
```
### 2.4 MCP + A2A: The Dual Foundation
```
┌────────────────────────────────────────────────────┐
│ AGENT ORCHESTRATION LAYER │
│ ┌──────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Planner │ │ State/Known │ │ Quality/Ops │ │
│ │ (delegate│ │ (memory) │ │ (monitor) │ │
│ │ tasks) │ │ │ │ │ │
│ └────┬─────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │
│ ┌────▼───────────────▼──────────────────▼───────┐ │
│ │ AGENT COMMUNICATION LAYER │ │
│ │ ┌─────────────────┐ ┌──────────────────┐ │ │
│ │ │ MCP (tools) │ │ A2A (peers) │ │ │
│ │ │ client-server │ │ peer-to-peer │ │ │
│ │ │ tool calls │ │ task delegation │ │ │
│ │ │ data access │ │ capability disc │ │ │
│ │ └─────────────────┘ └──────────────────┘ │ │
│ └───────────────────────────────────────────────┘ │
└────────────────────────────────────────────────────┘
MCP answers: "How does this agent get tools?"
A2A answers: "How does this agent talk to other agents?"
```
---
## 3. Orchestration Patterns Catalog
### 3.1 Sequential Pipeline (Level 1)
```
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐
│ Agent A │ ──► │ Agent B │ ──► │ Agent C │ ──► │ Agent D │
│ Parse │ │ Analyze │ │ Generate │ │ Validate │
└──────────┘ └──────────┘ └──────────┘ └──────────┘
│ │ │ │
└─────── Shared State ────────────┴────────────────┘
When to use: Fixed sequence, clear dependencies, batch processing
Cost profile: N agents, each sees previous output
Failure mode: Error cascades — bad output in step 1 poisons everything
No backtracking. Mitigation: checkpoint per step
Debugging: Trivial — failure at step N means error in step N
```
**Red team application:** Recon pipeline — SpiderFoot → Amass → Nmap → Vuln scan.
### 3.2 Supervisor / Orchestrator-Worker (Level 2)
```
┌──────────────┐
│ SUPERVISOR │
│ Plans, │
│ delegates, │
│ synthesizes │
└──┬───────┬───┘
┌──────┤ ├──────┐
│ │ │ │
┌──────▼──┐┌──▼────┐┌─▼─────┐
│ Worker 1││Worker2││Worker3│
│ Recon ││Exfil ││Report │
└─────────┘└───────┘└───────┘
When to use: Task decomposes cleanly, specialist agents
Need single point of accountability
Cost profile: 1 planner + N workers + 1 synthesizer
3-5 workers ideal. Beyond that, context overflow.
Failure mode: Supervisor is single point of failure.
Misclassification → wrong worker gets task.
Mitigation: Supervisor uses capable model, workers use cheaper ones.
Cost 40-60% less than running all on capable models.
```
**Red team application:** Campaign planning — supervisor defines TTPs, workers execute in parallel (phish, exploit, recon, pivot), synthesizer builds report.
### 3.3 Concurrent / Fan-Out / Fan-In (Level 2)
```
┌──────────────┐
│ DISPATCHER │
│ (same input) │
└──┬───┬───┬───┘
│ │ │
┌────────▼┐ ┌▼──┐ ┌▼────────┐
│ Agent A │ │ B │ │ Agent C │
│ Web │ │API│ │ DB │
└─────────┘ └───┘ └─────────┘
│ │ │
┌──▼───▼───▼──┐
│ COLLECTOR │
│ Vote / Merge│
│ / Synthesize│
└─────────────┘
When to use: Independent perspectives on same problem
Latency-critical: wall clock = max(agent)
Cost profile: N agents (parallel) + 1 aggregator
Failure mode: API rate limits (N agents × requests > limit)
Race conditions: N agents have N(N-1)/2 potential conflicts
LLM-based synthesis can hallucinate false consensus
Mitigation: Use explicit voting/weighted merge for deterministic aggregation
Cap at 5 parallel agents per dispatcher
```
**Red team application:** Multi-source data collection — SpiderFoot + Shodan + cert.sh + Amass simultaneously.
### 3.4 Handoff (Level 3)
```
┌─────────┐ ┌─────────┐ ┌─────────┐
│ Triage │ ──► │ Recon │ ──► │ Exploit │
│ Agent │ │ Agent │ │ Agent │
│ (class.)│ │ (enumer)│ │ (breach)│
└─────────┘ └─────────┘ └─────────┘
│
│ Only one active at a time
│ Full control transfers with context
When to use: Optimal specialist emerges during processing
Cheapest multi-agent pattern
Cost profile: 1 active agent × chain depth
Failure mode: INFINITE HANDOFF LOOPS — #1 production failure
Context loss compounds with each transfer
Mitigation: Max handoff depth (3-5). Explicit handoff schema (JSON).
Transfer only structured context, not full conversation.
```
**Red team application:** Triage → intelligence gathering → exploit chain. Each stage hands off to next specialist.
### 3.5 Group Chat / Debate (Level 3)
```
┌─────────────────────────────────────────────────┐
│ SHARED CONVERSATION THREAD │
│ │
│ Agent A: "I found X. Recommend approach Y." │
│ Agent B: "X is a false positive. Check Z." │
│ Agent C: "Confirmed. Z has a known CVE." │
│ Agent A: "Agreed. Path is Z via CVE-2026-XXXX." │
│ │
│ Chat Manager controls: who speaks, when to stop │
└─────────────────────────────────────────────────┘
When to use: Quality verification, maker-checker loops
Requires multi-perspective validation
Cost profile: N agents × rounds
Cap at 3 agents — beyond that, sycophancy cascades
Failure mode: Conversation loops (agents never converge)
Sycophancy: agents agree with majority even when wrong
Mitigation: Set max rounds (3-5). Maker uses cheap model,
checker uses capable model. Cost savings: 40-60%.
```
**Red team application:** Cross-validation of recon findings. One agent generates attack path, another validates, third checks for detection gaps.
### 3.6 Magentic / Adaptive Planning (Level 4)
```
┌─────────────────────────────────────────────────────┐
│ MANAGER AGENT │
│ Builds TASK LEDGER (live, revised) │
│ │
│ [ ] Subgoal 1: Recon network │
│ [✓] Subgoal 2: Identify entry points │
│ [ ] Subgoal 3: Craft exploit payload (reassigned to Agent X) │
│ [ ] Subgoal 4: Deploy persistence │
│ [ ] Subgoal 5: Exfil target data │
│ │
│ Manager reorders, reassigns, adds as context evolves │
└─────────────────────────────────────────────────────┘
When to use: Open-ended problems, no known solution path
Must discover approach during execution
Cost profile: Unbounded until plan converges
Failure mode: Slow to converge. Stalls on ambiguous goals.
Hard to estimate cost/time upfront.
Mitigation: Set max iteration budget (token/cost cap).
After Nth iteration, escalate to human.
```
**Red team application:** Complex multi-stage campaigns where path depends on intermediate discoveries.
### 3.7 Blackboard (Level 4)
```
┌─────────────────────────────────────────────┐
│ BLACKBOARD (shared state) │
│ │
│ key: "target.internal.ip" │
│ val: "10.0.1.55" │
│ key: "vulnerabilities" │
│ val: ["CVE-2026-XXXX on port 443"] │
│ key: "credentials" │
│ val: {"user": "admin", "hash": "..."} │
│ key: "task_status" │
│ val: {"recon": "done", "exploit": "fail" } │
└─────────────────────────────────────────────┘
▲ ▲ ▲ ▲
│ │ │ │
┌─────┴──┐ ┌─────┴──┐ ┌─────┴──┐ ┌─────┴──┐
│ Recon │ │Exploit │ │Pivot │ │Exfil │
│ Agent │ │Agent │ │Agent │ │Agent │
└────────┘ └────────┘ └────────┘ └────────┘
When to use: Dynamic, event-driven collaboration
Next step depends on intermediate findings
Failure mode: Write conflicts (two agents update same key)
Infinite loops (A writes → B triggers → A triggers)
Mitigation: Optimistic locking (Redis). Max iteration counter per item.
Task-completion flags prevent re-processing.
```
**Red team application:** Multi-phase op where each agent writes findings to shared board and other agents react. Recon writes a discovered port → Exploit agent reads it and attacks → writes result → Pivot agent reads and moves laterally.
### 3.8 Swarm (Level 5)
```
┌──────────────────────────────────────────────────┐
│ No central orchestrator. Agents route to peers. │
│ │
│ ┌─────────┐ │
│ │ Agent A │◄────► Agent B │
│ └────┬────┘ └────┬────┘ │
│ │ │ │
│ ┌────▼────┐ ┌────▼────┐ │
│ │ Agent C │◄───►│ Agent D │ │
│ └─────────┘ └─────────┘ │
│ │
│ Each agent knows its own capability set │
│ Each agent routes to peer when out of depth │
└──────────────────────────────────────────────────┘
When to use: Highly unpredictable tasks, emergent problem-solving
Failure mode: Near-impossible to debug. Non-deterministic routing.
Hard to predict cost or path.
Mitigation: Not for production systems without extreme observability.
Use only for research/exploration.
```
**Red team application:** Experimental — autonomous hunt teams where recon agents discover surfaces and route to exploit agents dynamically.
---
## 4. Security Architecture
### 4.1 The Agentic Security Stack
```
┌─────────────────────────────────────────────────────────────┐
│ LAYER 5: GOVERNANCE & POLICY │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ RBAC · ABAC · OPA policies · Data classification │ │
│ │ Per-workflow token budgets · Cost limits │ │
│ │ Compliance evidence (SOC 2, HIPAA, PCI) │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 4: ORCHESTRATION SECURITY │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Supervisor agent hardening · Planner constraints │ │
│ │ Delegation depth limits · Max iteration budgets │ │
│ │ Human-in-the-loop gates for destructive actions │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 3: AGENT IDENTITY & AUTHENTICATION │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ AIP (Agent Identity Protocol) — cryptographic identity │ │
│ │ Delegation tokens (Biscuit/macaroon) — attenuate scope │ │
│ │ OAuth 2.1 + PKCE — every agent has a verifiable ID │ │
│ │ Token: actor + subject claims, scope narrows per hop │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 2: MCP SECURITY │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ MCP Gateway — central authN/Z, policy, rate limits │ │
│ │ STDIO hardening — minimal container, non-root, R/O FS │ │
│ │ Supply chain — version pinning, Sigstore verification │ │
│ │ Tool poisoning detection — audit tool descriptions │ │
│ └─────────────────────────────────────────────────────────┘ │
├─────────────────────────────────────────────────────────────┤
│ LAYER 1: EXECUTION SANDBOX │
│ ┌─────────────────────────────────────────────────────────┐ │
│ │ Firecracker / gVisor micro-VMs for agent execution │ │
│ │ seccomp / SELinux profiles per agent type │ │
│ │ Network egress allowlists — deny all by default │ │
│ │ Read-only filesystem for all non-writing agents │ │
│ └─────────────────────────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────┘
```
### 4.2 The "Config is Code" Principle (from Paperclip CVE)
The foundational security insight from CVE-2026-41679: **agent configuration IS executable code.**
```
Every one of these must be treated as code:
├── YAML/JSON agent definitions
├── MCP server STDIO command strings
├── Tool descriptions (name, schema, instructions)
├── Prompt templates
├── Parameter schemas
├── Agent skill files (SKILL.md)
├── Orchestration workflow definitions
├── Inter-agent handoff schemas
└── Role/backstory prompts
Security posture:
- Code review for every config change
- Version control (git) for all definitions
- Sign config artifacts (Sigstore/Cosign)
- Pin versions — never deploy "latest"
- Audit trail on every config modification
```
### 4.3 MCP Security Hardening (Expanded)
From the Coalition for Secure AI RSAC 2026 findings, Knostic scan of ~2,000 MCP servers (all lacked auth), and the Microsoft MCP Security 2026 report:
```
┌────────────────────┬─────────────────────────────────────────┐
│ Risk │ Mitigation │
├────────────────────┼─────────────────────────────────────────┤
│ Prompt injection │ Validate all tool descriptions + params │
│ (tool poisoning) │ Separate trusted/untrusted MCP servers │
│ │ Never craft prompts from user content │
├────────────────────┼─────────────────────────────────────────┤
│ No auth on MCP │ OAuth 2.1 + PKCE per server │
│ servers │ No long-lived tokens. Rotate every hour. │
│ (Knostic finding: │ Token audience bound to specific server │
│ ~2,000 servers, │ │
│ 0 with auth) │ │
├────────────────────┼─────────────────────────────────────────┤
│ Supply chain rug │ Pin versions. Verify build provenance. │
│ pull │ Watch for ownership changes on upstream │
│ │ SBOM per MCP server │
├────────────────────┼─────────────────────────────────────────┤
│ Shadow MCP │ Inventory all MCP servers. Unregistered │
│ servers │ = blocked by default at gateway. │
│ │ Auto-discover via network scan. │
├────────────────────┼─────────────────────────────────────────┤
│ Over-permissioned │ Narrow OAuth scopes. "Delete" scope only │
│ tokens │ when agent needs to delete. Audited. │
│ │ Scoped per resource, not per server. │
├────────────────────┼─────────────────────────────────────────┤
│ Multi-hop identity │ AIP (Agent Identity Protocol). │
│ chain loss │ Delegation tokens with attenuated scope. │
│ │ Policy evaluated at every hop. │
└────────────────────┴─────────────────────────────────────────┘
```
### 4.4 The Blast Radius Principle
```
┌───────────────────────────────────────────────────────────┐
│ BLAST RADIUS = every agent can reach │
│ ├── These MCP servers │
│ ├── These tools │
│ ├── This memory/state │
│ ├── These peer agents │
│ └── These filesystem paths │
│ │
│ RULE: Narrow the blast radius of every agent to the │
│ absolute minimum required for its role. │
│ │
│ Example: Recon agent needs: │
│ ✓ web_search() tool │
│ ✗ write_file() tool │
│ ✗ db_query() tool │
│ ✗ access to peer agents │
│ ✗ access to credential store │
│ │
│ Exploit agent needs: │
│ ✓ command execution (sandboxed) │
│ ✓ write_file() to output directory │
│ ✓ A2A comms with Pivot agent │
│ ✗ access to production database │
│ ✗ ability to exfiltrate to arbitrary endpoints │
└───────────────────────────────────────────────────────────┘
```
---
## 5. State & Memory Architecture
### 5.1 Three-Tier Memory Model
```
┌────────────────────────────────────────────────────────────┐
│ TIER 1: WORKING MEMORY (ephemeral, per-task) │
│ Storage: In-context (LLM context window) │
│ Lifetime: Single agent call │
│ Contents: Current reasoning, tool outputs, partial state │
│ Risk: Context window overflow │
│ Mitigation: Token budgets, message compression │
├────────────────────────────────────────────────────────────┤
│ TIER 2: SESSION MEMORY (persistent, per-conversation) │
│ Storage: State store (Redis, SQLite, file) │
│ Lifetime: Entire agent session (hours to days) │
│ Contents: Conversation history, completed tasks, │
│ intermediate findings, credentials cache │
│ Key property: Survives agent restarts, context pruning │
│ Mitigation: Session expiry, max size limits, PII redact │
├────────────────────────────────────────────────────────────┤
│ TIER 3: SEMANTIC MEMORY (long-term, cross-session) │
│ Storage: Vector DB / knowledge graph (Chroma, pgvector) │
│ Lifetime: Months to permanent │
│ Contents: Learned facts, reusable procedures, │
│ skill definitions, absorbed knowledge │
│ Key property: Injected into new sessions as context │
│ Mitigation: Trust metadata, provenance tracking, │
│ periodic consolidation, corruption detection │
└────────────────────────────────────────────────────────────┘
```
### 5.2 Orchestration State Management
```
┌──────────────────────────────────────────────────────────┐
│ ORCHESTRATOR STATE (per workflow execution) │
│ │
│ { │
│ "workflow_id": "recon-20261003-001", │
│ "status": "in_progress", │
│ "plan": { │
│ "phases": ["recon", "exploit", "exfil"], │
│ "current_phase": "recon", │
│ "agents_assigned": { │
│ "recon": ["spiderfoot_mcp", "amass_agent"], │
│ "exploit": null, │
│ "exfil": null │
│ } │
│ }, │
│ "artifacts": { │
│ "domains_found": ["target.com", "admin.target.com"],│
│ "ports_open": [80, 443, 8443], │
│ "cves_possible": ["CVE-2026-XXXX"] │
│ }, │
│ "errors": [], │
│ "checkpoints": [ │
│ {"phase": "recon", "time": "2026-10-03T02:00:00Z", │
│ "state_hash": "abc123"} │
│ ] │
│ } │
│ │
│ Key properties: │
│ - Checkpointed at every phase (survives failure) │
│ - State hash for integrity verification │
│ - Error log with rollback capability │
│ - Artifacts accumulate across phases │
└──────────────────────────────────────────────────────────┘
```
### 5.3 Inter-Agent Context Handoff
When agents hand off tasks, pass **structured schema, not conversation history:**
```
BAD: Full conversation dump (context expensive, LLM distills lossily)
Agent A → Agent B: [50K tokens of raw chat history]
GOOD: Structured handoff object (schema-defined, compressed, lossless)
Agent A → Agent B: {
"task": "exploit_port_443",
"findings": {
"target_ip": "10.0.1.55",
"port": 443,
"service": "nginx 1.24.0",
"cve": "CVE-2026-XXXX",
"confidence": 0.85,
"source": "nmap-agent"
},
"context_chain": ["recon", "scan", "identify"],
"constraints": {
"max_duration_s": 300,
"destructive_allowed": false
}
}
```
---
## 6. Framework Selection Guide
### 6.1 Comparison Matrix
```
┌──────────────────┬───────────┬───────────┬────────────┬───────────┬──────────────┐
│ Framework │ Best For │ Orchestr. │ State Mgmt │ Security │ MCP Support │
│ │ │ Model │ │ First? │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ LangGraph │ Complex │ DAG / │ Checkpoint │ No (add │ Full (native)│
│ │ stateful │ State │ per node, │ OPA/ │ │
│ │ workflows │ Graph │ time-travel│ gateway) │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ CrewAI │ Role- │ Crews / │ Task-based │ No (add │ Full │
│ │ based │ Hierarch. │ sequential │ gateway) │ │
│ │ business │ │ memory │ │ │
│ │ processes │ │ │ │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ MS Agent │ Enterprise│ Graph / │ Session │ Yes │ Full (native)│
│ Framework │ Azure │ GroupChat │ state + │ (RBAC, │ │
│ │ stack │ Handoff │ checkpoint │ OAuth) │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ Google ADK │ Google │ Hierarch. │ Session │ Partial │ Full + A2A │
│ │ Cloud │ tree + │ (per agent)│ │ (donated to │
│ │ env │ A2A │ │ │ LF) │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ OpenAI Agents │ Tuned to │ Handoff │ External │ No (add │ Full │
│ SDK │ OpenAI │ (tool │ (BYO DB) │ gateway) │ │
│ │ models │ call) │ │ │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ PyAgent │ 18 │ All 4 │ 3-tier │ Yes (OPA)│ Full │
│ │ patterns │ tiers │ memory │ │ │
│ │ out-of-box│ │ │ │ │
├──────────────────┼───────────┼───────────┼────────────┼───────────┼──────────────┤
│ Hermes Agent │ CLI agent │ None │ Memory / │ Limited │ Full │
│ (this runtime) │ lifecycle │ (handoff │ Session │ (profiles)│ │
│ │ │ via │ search │ │ │
│ │ │ delegate) │ │ │ │
└──────────────────┴───────────┴───────────┴────────────┴───────────┴──────────────┘
```
### 6.2 Selection Decision Tree
```
┌──────────────────────────────────────────────────────────────┐
│ Are you on a specific cloud stack? │
│ Microsoft / Azure → MS Agent Framework │
│ Google Cloud → Google ADK │
│ AWS / self-hosted → LangGraph or PyAgent │
│ None / multi-cloud → LangGraph or CrewAI │
├──────────────────────────────────────────────────────────────┤
│ Is security-first governance required? │
│ YES → MS Agent Framework (Entra ID) or PyAgent (OPA built-in)│
│ Add MCP Gateway layer regardless of framework choice │
├──────────────────────────────────────────────────────────────┤
│ Quick prototype needed? │
│ Role-based → CrewAI (fastest to working multi-agent) │
│ Tool-heavy → OpenAI Agents SDK (minimal abstraction) │
├──────────────────────────────────────────────────────────────┤
│ Production reliability critical? │
│ Deterministic DAG → LangGraph (checkpointing, time-travel) │
│ Complex routing → Conductor (Microsoft, Jinja2 routing) │
├──────────────────────────────────────────────────────────────┘
```
---
## 7. Red Team Integration Mapping
### 7.1 Pattern-to-Kill-Chain Mapping
```
┌────────────────────────────┬─────────────────────────────────┐
│ Red Team Phase │ Optimal Orchestration Pattern │
├────────────────────────────┼─────────────────────────────────┤
│ Reconnaissance (TA0043) │ Concurrent / Fan-Out │
│ │ SpiderFoot + Amass + Shodan │
│ │ in parallel → collector merges │
├────────────────────────────┼─────────────────────────────────┤
│ Resource Dev (TA0042) │ Sequential pipeline │
│ │ Domain buy → DNS setup → │
│ │ cert → VPS provision │
├────────────────────────────┼─────────────────────────────────┤
│ Initial Access (TA0001) │ Handoff │
│ │ Triage → phish → exploit chain │
├────────────────────────────┼─────────────────────────────────┤
│ Execution (TA0002) │ Supervisor-worker │
│ │ Deploy dropper, stager, payload │
├────────────────────────────┼─────────────────────────────────┤
│ Persistence (TA0003) │ Blackboard │
│ │ Agents monitor and react to │
│ │ access state changes │
├────────────────────────────┼─────────────────────────────────┤
│ PrivEsc (TA0004) │ Handoff │
│ │ Enum → attempt → next technique │
├────────────────────────────┼─────────────────────────────────┤
│ Credential Access (TA0006) │ Concurrent │
│ │ LSASS + browser + Kerberoast │
│ │ all parallel → merge findings │
├────────────────────────────┼─────────────────────────────────┤
│ Lateral Movement (TA0008) │ Blackboard │
│ │ Each agent writes pivot points │
│ │ Others read and exploit │
├────────────────────────────┼─────────────────────────────────┤
│ Exfiltration (TA0010) │ Sequential pipeline │
│ │ Collect → archive → encrypt → │
│ │ exfil → verify │
├────────────────────────────┼─────────────────────────────────┤
│ Reporting │ Supervisor-worker │
│ │ Planner assigns sections to │
│ │ specialist writer agents │
└────────────────────────────┴─────────────────────────────────┘
```
### 7.2 Agent Roles in a Red Team MAS
```
┌────────────────┬────────────────────────────┬─────────────────┐
│ Agent Role │ Responsibilities │ MCP Tools │
├────────────────┼────────────────────────────┼─────────────────┤
│ Recon Agent │ OSINT, surface mapping, │ web_search, │
│ │ subdomain enumeration, │ web_extract, │
│ │ technology profiling │ SpiderFoot API │
├────────────────┼────────────────────────────┼─────────────────┤
│ Intrusion Agent│ Phishing, exploit dev, │ SMTP relay, │
│ │ payload generation, │ payload builder,│
│ │ initial access │ browser │
├────────────────┼────────────────────────────┼─────────────────┤
│ C2 Agent │ Beacon deployment, │ Sliver/Mythic │
│ │ heartbeat management, │ API, MCP -> C2 │
│ │ command dispatch │ bridge │
├────────────────┼────────────────────────────┼─────────────────┤
│ Pivot Agent │ Lateral movement, │ impacket, SSH, │
│ │ AD exploitation, cloud │ BloodHound API │
│ │ escalation │ │
├────────────────┼────────────────────────────┼─────────────────┤
│ Intel Agent │ Credential parsing, │ Mimikatz API, │
│ │ data collection, │ hash parser, │
│ │ intelligence aggregation │ data archiver │
├────────────────┼────────────────────────────┼─────────────────┤
│ Report Agent │ Findings aggregation, │ Markdown writer,│
│ │ ATT&CK mapping, │ ATT&CK Nav API, │
│ │ report generation │ template engine │
├────────────────┼────────────────────────────┼─────────────────┤
│ Watchdog Agent │ OPSEC monitoring, │ log parser, │
│ │ detection checks, │ network scanner,│
│ │ kill switch trigger │ alert handler │
└────────────────┴────────────────────────────┴─────────────────┘
```
### 7.3 Novel: MCP Bridge for C2 Operations
Building on SANS SEC565 (2026) research — using MCP servers as a C2 transport layer:
```
┌─────────────────────────────────────────────────────────────┐
│ MCP Bridge C2 Architecture │
│ │
│ ATTACKER SIDE TARGET SIDE │
│ ┌────────────────┐ ┌──────────────────┐ │
│ │ C2 Orchestrator│ │ MCP Agent Runtime│ │
│ │ Agent │ │ │ │
│ │ │ MCP STDIO │ "tools/call" │ │
│ │ execute cmd ───┼─────────────► execute(name: │ │
│ │ │ │ "shell_cmd") │ │
│ │ │◄─────────────┤ return: output │ │
│ │ read output │ │ │ │
│ └────────────────┘ └──────────────────┘ │
│ │
│ The MCP STDIO transport is indistinguishable from │
│ legitimate agent tool usage to network monitoring. │
│ C2 commands are JSON-RPC "tools/call" messages. │
│ │
│ RISK: This is novel (2026). Detection coverage is low. │
│ ADVANTAGE: Blends with normal agent traffic on the host. │
│ DEFENSE: Monitor for suspicious tool definitions in MCP │
│ server configs. Audit all STDIO server commands. │
└─────────────────────────────────────────────────────────────┘
```
---
## 8. Observability & Audit
### 8.1 What to Trace
Every agent interaction must produce structured logs:
```
┌─────────────────────────────────────────────────────────────┐
│ AGENT TRACE RECORD │
│ { │
│ "trace_id": "recon-20261003-001", │
│ "span_id": "spiderfoot-mcp-call-003", │
│ "parent_span": "supervisor-plan-001", │
│ "agent": "recon-agent-v1", │
│ "action": "tools/call", │
│ "tool": "web_search", │
│ "input_hash": "sha256:abc...", │
│ "output_hash": "sha256:def...", │
│ "tokens_in": 452, │
│ "tokens_out": 128, │
│ "cost_usd": 0.0042, │
│ "duration_ms": 3402, │
│ "status": "success", │
│ "timestamp": "2026-10-03T02:00:00Z", │
│ "model": "claude-opus-4.6-1m" │
│ } │
│ │
│ Collect into: OpenTelemetry-compatible trace store │
│ Visualize in: LangSmith, Weights & Biases, custom dashboard │
│ Query via: trace_id across spans → full execution timeline │
└─────────────────────────────────────────────────────────────┘
```
### 8.2 Cost Tracking Per Pattern
```
Set per-workflow budgets BEFORE execution:
Sequential (4 agents): Max 50K tokens / $0.15
Supervisor (1+3): Max 80K tokens / $0.25
Concurrent (4 agents): Max 60K tokens / $0.18
Handoff (3 hops): Max 40K tokens / $0.12
Group chat (3×3 rounds): Max 90K tokens / $0.28
Magentic: CIRCUIT BREAKER at 200K tokens
Gateways enforce budgets. Exceeded = circuit breaker + notification.
```
### 8.3 Audit Trail Requirements
```
Non-negotiable for red team agent ops:
[✓] Every tool call logged with input_hash and output_hash
[✓] Every agent handoff logged with full context summary
[✓] Every MCP server connection logged (source agent, server, transport)
[✓] Every configuration change logged (who, what, when, diff)
[✓] Every destructive action requires HITL gate + audit record
[✓] Token/cost budget per workflow, per agent, per session
[✓] Trace ID chains: user request → orchestrator → agent → MCP call
```
---
## 9. Deployment Patterns
### 9.1 Deterministic vs Dynamic Orchestration
```
┌──────────────────────┬──────────────────────────┬────────────────┐
│ │ DETERMINISTIC (DAG) │ DYNAMIC (LLM │
│ │ │ routing) │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Topology │ Fixed at design time │ Emerges at │
│ │ │ runtime │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Debugging │ Trivial — step N fails │ Hard — non- │
│ │ = agent N has the bug │ deterministic │
│ │ │ execution path │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Cost │ Predictable (N × tokens) │ Variable (may │
│ │ │ loop/explode) │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Best for │ Known workflows, batch │ Discovery, │
│ │ processing, ETL │ research, open │
│ │ │ problems │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Blast radius │ Bounded by design │ Uncapped │
├──────────────────────┼──────────────────────────┼────────────────┤
│ Audit ability │ Full traceability │ Partial (LLM │
│ │ │ decision not │
│ │ │ deterministic) │
└──────────────────────┴──────────────────────────┴────────────────┘
RECOMMENDATION: Default to DAG. Reserve dynamic for the specific
sub-problem that genuinely requires it.
```
### 9.2 Hybrid Patterns (Production Norm)
Production systems rarely use one pure pattern. Common hybrid:
```
┌──────────────────────────────────────────────────────────┐
│ SUPERVISOR at top level │
│ │ │
│ ├── PIPELINE for stage 1 (recon: fixed sequence) │
│ │ SpiderFoot → Amass → Nmap → Vuln scan │
│ │ │
│ ├── CONCURRENT for stage 2 (multi-source collection) │
│ │ cert.sh + Shodan + GitHub dorking in parallel │
│ │ │
│ ├── HANDOFF for stage 3 (triage → exploit chain) │
│ │ Triage agent → Exploit agent → Pivot agent │
│ │ │
│ └── SUPERVISOR for stage 4 (report synthesis) │
│ Writers generate sections in parallel │
│ Supervisor assembles final report │
│ │
│ Each stage gets its own pattern. │
│ The supervisor decides which stage to run next. │
└──────────────────────────────────────────────────────────┘
```
---
## 10. Hardening Checklist
### 10.1 Pre-Deployment Security Gates
```
[ ] Every agent definition code-reviewed
[ ] Every MCP server version-pinned + provenance verified
[ ] Every STDIO server command string audited (is it executable?)
[ ] OAuth 2.1 + PKCE configured on every MCP server
[ ] Tokens scoped per-resource (never wildcard)
[ ] Token rotation policy set (max 1 hour for operational agents)
[ ] Gateway in path — no agent talks directly to MCP server
[ ] Network egress deny-by-default for sandboxed agents
[ ] Read-only filesystem for all non-writing agents
[ ] Per-workflow token/cost budgets defined
[ ] Circuit breakers configured (max iterations, max cost)
[ ] Human-in-the-loop gates for destructive actions
[ ] All agent interactions traced to OpenTelemetry store
[ ] Trace IDs chainable from user request to individual tool call
[ ] Shadow MCP detection active (inventory scan)
[ ] Supply chain monitoring on all MCP server dependencies
[ ] Agent identity attested (AIP or similar)
[ ] Delegation scope narrows at every hop (never expands)
```
### 10.2 Runtime Monitoring
```
┌──────────────────────────┬──────────────────────────────┐
│ What to Monitor │ Alert Threshold │
├──────────────────────────┼──────────────────────────────┤
│ Orchestration loops │ >5 iterations without progress│
│ Agent handoff depth │ >5 hops without completion │
│ Token consumption/workfl.│ >2x estimated budget │
│ Cost/workflow │ >$0.50 USD │
│ MCP server response │ >30s timeout │
│ Agent context window │ >90% of limit │
│ Inter-agent latency │ >10s between handoffs │
│ Tool call failure rate │ >10% of calls │
│ Unregistered MCP servers │ Any = immediate │
│ Long-lived tokens │ >1 hour without rotation │
│ Unattenuated delegation │ Scope non-narrowing at hop │
└──────────────────────────┴──────────────────────────────┘
```
---
## 11. Tool & Resource Reference
### 11.1 Frameworks (2026)
| Framework | URL | License | Key Strength |
|-----------|-----|---------|--------------|
| LangGraph | github.com/langchain-ai/langgraph | MIT | Stateful DAG, checkpointing |
| CrewAI | github.com/crewAIInc/crewAI | MIT | Fast role-based prototype |
| MS Agent Framework | github.com/microsoft/agent-framework | MIT | Enterprise, Azure, security |
| Google ADK | github.com/google/adk | Apache 2.0 | A2A protocol, GCP native |
| OpenAI Agents SDK | github.com/openai/openai-agents-python | MIT | Clean handoff, MCP native |
| PyAgent | github.com/pyagent-core/pyagent | MIT | 18 patterns, OPA built-in |
| Conductor | opensource.microsoft.com | MIT | Deterministic, Jinja2 routing |
### 11.2 MCP Ecosystem
| Tool | Purpose | Security Note |
|------|---------|---------------|
| MCP Gateway | Centralized authN/Z, policy, audit | REQUIRED for production |
| ScaleMCP | Dynamic tool inventory sync | Needs gateway integration |
| AgentMaster | MCP + A2A integration comoposer | Experimental |
| MCP Inspector | Debug/audit MCP server configs | Scan for STDIO injection |
| MCP OWASP Top 10 | Risk framework for MCP | Maps to Azure/OPA controls |
### 11.3 Observability
| Tool | Purpose |
|------|---------|
| LangSmith | Framework-agnostic tracing, eval |
| OpenTelemetry | Standard collector and export |
| Weights & Biases | Experiment tracking, cost analysis |
| VECTR | Red team engagement tracking |
| ATT&CK Navigator | Technique coverage visualization |
---
## 12. The Paperclip Lesson (Reinforced)
CVE-2026-41679 taught us: **"Configuration is code."** This principle applies at every level of agent orchestration:
```
Agent YAML/JSON config → IS CODE
MCP server STDIO command → IS CODE
Tool description schema → IS CODE (tool poisoning)
Inter-agent handoff schema → IS CODE
Orchestration workflow def → IS CODE
Skill file (SKILL.md) → IS CODE
Prompt template → IS CODE
Role/backstory assignment → IS CODE (prompt injection)
Every one of these can be exploited.
Every one of these must be reviewed, signed, pinned, and audited.
```
**The attack chain that works against any orchestrated system:**
```
1. Attacker submits crafted input to an agent
2. Agent treats input as data — but orchestrator treats it as workflow
3. Input becomes plan → determines which agent runs, with what tools
4. Orchestrator delegates to a privileged agent
5. Privileged agent executes attacker's intended action
⚠ This is not theoretical. This is Paperclip applied to orchestration.
⚠ Mitigation: Validate at every hop. Never trust inputs as workflow.
```
---
**End of ptSlick Agentic Orchestrated Blueprint v1.0**
> "Start with a single agent. Orchestrate only when the limitation forces it.
> Monitor every handoff. Audit every config. Budget every token."