PrimeSwarm vs Microsoft Agent Governance Toolkit
An honest comparison with the toolkit that covers both layers — where AGT is broader, where PrimeSwarm is deeper, where they could be complementary
PrimeSwarm vs Microsoft Agent Governance Toolkit
An honest comparison with a toolkit that covers both layers
Microsoft released the Agent Governance Toolkit (AGT) as an MIT-licensed, public-preview toolkit for governing autonomous AI agents. It covers policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering. It maps to OWASP Agentic Top 10 (7/10 full, 3/10 partial), NIST AI RMF, EU AI Act, and SOC 2. It has 992 conformance tests, 29 architecture decision records, 10 formal specifications, 14 framework integrations, and 5 language SDKs.
This is the project that comes closest to covering both what PrimeSwarm does and what NVIDIA NeMoClaw does — in one toolkit. It deserves an honest comparison.
This paper does not claim PrimeSwarm wins. It states where AGT is stronger, where PrimeSwarm is stronger, and where the two could be complementary. If you are a security architect evaluating both, this is what we think you should know.
What AGT is
AGT is a 9-package Python toolkit with TypeScript, .NET, Rust, and Go SDKs:
| Package | What it does |
|---|---|
| Agent OS | Policy engine, agent lifecycle, governance gate |
| Agent Control Specification (ACS) | Stateless, deterministic, fail-closed policy runtime with a Rust core |
| Agent Mesh | Agent discovery, routing, trust mesh, Ed25519 identity, DID-based credentials |
| Agent Runtime | Execution sandboxing with 4 privilege rings (Ring 0-3), session isolation, kill switch |
| Agent SRE | Circuit breakers, error budgets, chaos testing, cascade detection |
| Agent Compliance | OWASP verification, policy linting, integrity checks |
| Agent Marketplace | Plugin governance and trust scoring |
| Agent Lightning | RL training governance with violation penalties |
| Agent Hypervisor | Execution audit, delta engine, in-memory commitment tracking, command denylist |
Additional capabilities include an MCP Security Gateway (tool poisoning detection, drift monitoring, typosquatting, hidden instruction scanning), Shadow AI Discovery (find unregistered agents), a Governance Dashboard, and a 12-vector PromptDefense Evaluator.
The policy format is YAML with Python-expression conditions:
rules:
- name: block-destructive
condition: "action.type in ['drop', 'delete', 'truncate']"
action: deny
The verdicts are allow, warn, deny, escalate, or transform. The runtime is stateless, deterministic, and fails closed. The audit trail is a Merkle hash chain — each entry contains the SHA-256 of the previous entry.
Where AGT is stronger than PrimeSwarm
We do not pretend otherwise. AGT is broader.
Sandboxing and execution isolation
AGT's Agent Runtime provides 4 execution rings (Ring 0-3) with privilege isolation, session isolation with filesystem-level boundaries, a kill switch with step handoff and compensation, and saga orchestration for multi-step transactions with automatic rollback. PrimeSwarm does not provide container isolation. We said this on our own threat model page — we govern decisions, not processes.
Multi-agent identity and trust
AGT's Agent Mesh provides Ed25519-backed identity, DID-style agent credentials, trust scoring with decay and revocation, challenge-response handshakes, and delegated trust chains. PrimeSwarm governs a single agent's decisions. We do not have a multi-agent mesh.
SRE and reliability engineering
AGT's Agent SRE provides circuit breakers, error budgets, SLO-driven enforcement, chaos testing, and cascade detection across agent fleets. PrimeSwarm does not have this.
Framework support
AGT integrates with 14 frameworks: Microsoft Agent Framework, Semantic Kernel, AutoGen, LangGraph, LangChain, CrewAI, OpenAI Agents SDK, Claude Code, Google ADK, LlamaIndex, Haystack, Mastra, Dify, and Azure AI Foundry. PrimeSwarm has limited framework integrations.
Language SDKs
AGT ships SDKs in 5 languages: Python, TypeScript, .NET, Rust, and Go. PrimeSwarm has Rust and TypeScript.
MCP security
AGT's MCP Security Gateway detects tool poisoning, drift, typosquatting, and hidden instructions in MCP servers. PrimeSwarm has an MCP tool protocol but no security gateway.
Compliance framework mapping
AGT maps to OWASP Agentic Top 10, NIST AI RMF, EU AI Act, SOC 2, AARM Extended, and the Agentic Trust Framework. PrimeSwarm has the DGV test card suite but no formal mapping to these frameworks.
Scale and backing
AGT is backed by Microsoft, has 992 conformance tests, 29 ADRs, an OpenSSF Scorecard, CodeQL, Gitleaks, ClusterFuzzLite, and Dependabot. PrimeSwarm is an independent research lab with 89 DGV test cards.
Where PrimeSwarm is stronger than AGT
This is where it matters. AGT's own documentation — specifically LIMITATIONS.md and the OWASP mapping — explicitly states gaps that PrimeSwarm addresses.
Knowledge governance (AGT Limitation #7)
AGT says, in their own words:
"AGT governs agent actions. It does not govern the knowledge agents consume — the documents, databases, embeddings, and context retrieved during reasoning. AGT does not verify the provenance, freshness, or authorization of documents retrieved via RAG. AGT does not track which knowledge sources influenced an agent's decision. AGT does not enforce data classification labels on retrieved context."
PrimeSwarm has provenance enforcement. No legal, clinical, or financial assertion without a source receipt. The Continuity Ledger records which knowledge sources influenced each decision. For regulated industries — where an agent citing the wrong regulation or a stale clinical guideline is a liability — this is the gap that matters.
Memory governance (AGT ASI06 — Partial)
AGT's OWASP mapping marks ASI06 (Memory and Context Poisoning) as Partial:
"No dedicated memory-sandbox or context-integrity module. The audit hash-chain provides tamper detection for any persisted state. However, AGT does not yet sandbox agent memory stores."
PrimeSwarm has the Continuity Ledger — durable governed memory with Ebbinghaus decay tuned to domain relevance. Memory is governed, not just audited. Clinical memory decays differently from financial memory. Recent lab values matter more than six-month-old notes. The Ledger enforces this.
Fairness and bias
AGT does not mention fairness gates, Adverse Impact Ratio, the Four-Fifths Rule, or protected-class analysis anywhere in their specifications, OWASP mapping, or limitations document. PrimeSwarm has fairness gates with automatic Adverse Impact Ratio calculation and Pre-Effect Refusal when bias is mathematically detected. For a lending platform, a hiring tool, or a clinical triage system, this is not optional. It is the law.
PII sanitization
AGT mentions "prompt/content sanitization" and "blocked_patterns to catch PII in outputs." PrimeSwarm has a dedicated PII sanitization layer that strips SSNs, account numbers, and NPI from context before the agent processes them, replacing them with cryptographic hashes. The agent never sees the raw PII. AGT catches PII on the way out. PrimeSwarm removes it on the way in.
ONLY Lang vs YAML conditions
AGT's policy format is YAML with Python-expression conditions. The condition "action.type in ['drop', 'delete', 'truncate']" is code. A compliance officer cannot read it. A developer must write it.
PrimeSwarm's ONLY Lang is a purpose-built governance DSL with mathematical semantics:
harmony([0.5, 0.3, 0.2])
assert_bounds(0, 50, 0)
budget_limit(500)
escalate
A compliance officer can read this. The policy is the policy — not code that implements the policy. For an auditor who needs to verify that the governance script matches the regulatory requirement, ONLY Lang is a different category of artifact.
Cryptographic receipts vs Merkle hash chain
AGT has a Merkle hash chain audit log. Each entry contains the SHA-256 of the previous entry, making tampering detectable. This is good. It answers: "what happened."
PrimeSwarm produces cryptographic receipts — SHA-256 of script + input hash + gate state + residual + timestamp, forwarded to SIEM, replayable with the open-source verifier. This answers: "replay this exact decision and verify it yourself."
The difference: an auditor reading an AGT audit log trusts that the log is accurate. An auditor reading a PrimeSwarm receipt does not need to trust the receipt — they can replay the evaluation with the open-source verifier and confirm the result independently. For an auditor who needs to verify a decision without trusting the platform that made it, receipts are stronger.
Regulated-industry deployment model
AGT is a toolkit. You pip install it and integrate it yourself. The documentation is excellent. The examples are thorough. But the deployment is your problem.
PrimeSwarm is a deployed service with a 4-phase engagement model:
- Discovery — we map your existing decision flow to the governance pipeline
- Build — we define the ONLY Lang scripts that encode your compliance rules
- Deploy — we deploy to your VPC, on-prem, or private cloud
- Maintain — quarterly governance reviews, DGV re-verification, ONLY Lang script updates for regulatory changes
For a regulated industry that needs a deployed service with ongoing governance maintenance — not a toolkit — this is the difference.
Where they are comparable
| Dimension | AGT | PrimeSwarm |
|---|---|---|
| Deterministic policy | ACS is stateless, deterministic, fail-closed | Gate is deterministic |
| Fail-closed | Yes — any error yields deny | Yes |
| Audit trail | Merkle hash chain | SHA-256 receipts + SIEM |
| HITL approval | Approval workflows with named approvers | HITL escalation with cryptographic sign-off |
| Open source | MIT | MIT/Apache (verifier, compiler, cards) |
| Prompt injection | Input rails, 12-vector PromptDefense Evaluator | TPNN spatial constraints |
The honest bottom line
AGT is the broader toolkit. If you want a general-purpose agent governance toolkit that integrates with 14 frameworks, runs in 5 languages, covers sandboxing, multi-agent identity, SRE, and 6 compliance frameworks — AGT is the stronger choice. We do not compete on breadth.
PrimeSwarm is deeper on the dimensions that matter to regulated industries. We fill gaps that AGT explicitly documents:
- Knowledge provenance (AGT Limitation #7)
- Memory governance (AGT ASI06 — Partial)
- Fairness and bias (not in AGT)
- PII sanitization before processing (AGT does it after)
- Cryptographic receipts vs hash-chain logs
- ONLY Lang as a compliance-readable DSL vs YAML conditions
- A deployed service vs a toolkit
They could be complementary. AGT could handle the infrastructure layer — sandboxing, identity, SRE, framework integration — while PrimeSwarm handles the decision layer — fairness, provenance, PII, receipts, memory governance. Both are deterministic. Both fail closed. Both produce audit trails. Running PrimeSwarm's ONLY Lang policies inside AGT's policy engine is architecturally feasible.
What we are not claiming
- We are not claiming PrimeSwarm is "better" than AGT. AGT is broader. We are narrower and deeper on specific dimensions.
- We are not claiming AGT is insufficient. AGT is excellent at what it does. Its limitations document is more honest than most vendors' marketing.
- We are not claiming ONLY Lang is universally superior to YAML conditions. For a developer, YAML is faster. For a compliance officer, ONLY Lang is readable. Different audiences.
- We are not claiming receipts are universally superior to hash chains. For internal auditing, hash chains are sufficient. For external auditors who need independent verification, receipts are stronger.
- We are not claiming the DGV test cards are superior to AGT's 992 conformance tests. They serve different audiences. AGT's tests verify the toolkit works. DGV cards verify the governance works.
What we are claiming
PrimeSwarm addresses gaps that AGT explicitly documents. For regulated industries — finance, healthcare, legal, accounting — those gaps are not optional. A lending platform without fairness gates is a lawsuit. A clinical agent without provenance enforcement is a malpractice risk. A legal agent without memory governance is a privilege violation.
If you are evaluating both, the question is not which is broader. The question is which gaps matter to your threat model.
Try it yourself
- Run the governance gate in your browser at /try — write an ONLY Lang script, set the field values, press Run, see the receipt.
- Browse all 89 DGV test cards at /dgv/registry — each one is a verifiable governance test.
- See the full threat model at /threat-model — what we cover, what we don't, what we recommend for the gaps.
- Read the AGT repository at github.com/microsoft/agent-governance-toolkit — it is excellent and we recommend you evaluate it honestly.
For the full verifier bundle — run all 89 cards locally and produce signed receipts — go to /try/bundle.
Published by Only Institute