Only Institute / Knowledge

Follow your curiosity.

AI sends your question and public excerpts to PrimeSwarm.

Try agent memory, governance, or a project name.

All Publications
whitepaper

PrimeSwarm vs Microsoft Agent Governance Toolkit

An honest comparison with the toolkit that covers both layers — where AGT is broader, where PrimeSwarm is deeper, where they could be complementary

Grigori Korotkikh 2026-09-12 14 min
PrimeSwarmMicrosoftAGTGovernanceCompetitionOWASP
Only Institute — PrimeSwarm vs Microsoft Agent Governance Toolkit

PrimeSwarm vs Microsoft Agent Governance Toolkit

An honest comparison with a toolkit that covers both layers

Microsoft released the Agent Governance Toolkit (AGT) as an MIT-licensed, public-preview toolkit for governing autonomous AI agents. It covers policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering. It maps to OWASP Agentic Top 10 (7/10 full, 3/10 partial), NIST AI RMF, EU AI Act, and SOC 2. It has 992 conformance tests, 29 architecture decision records, 10 formal specifications, 14 framework integrations, and 5 language SDKs.

This is the project that comes closest to covering both what PrimeSwarm does and what NVIDIA NeMoClaw does — in one toolkit. It deserves an honest comparison.

This paper does not claim PrimeSwarm wins. It states where AGT is stronger, where PrimeSwarm is stronger, and where the two could be complementary. If you are a security architect evaluating both, this is what we think you should know.

What AGT is

AGT is a 9-package Python toolkit with TypeScript, .NET, Rust, and Go SDKs:

PackageWhat it does
Agent OSPolicy engine, agent lifecycle, governance gate
Agent Control Specification (ACS)Stateless, deterministic, fail-closed policy runtime with a Rust core
Agent MeshAgent discovery, routing, trust mesh, Ed25519 identity, DID-based credentials
Agent RuntimeExecution sandboxing with 4 privilege rings (Ring 0-3), session isolation, kill switch
Agent SRECircuit breakers, error budgets, chaos testing, cascade detection
Agent ComplianceOWASP verification, policy linting, integrity checks
Agent MarketplacePlugin governance and trust scoring
Agent LightningRL training governance with violation penalties
Agent HypervisorExecution audit, delta engine, in-memory commitment tracking, command denylist

Additional capabilities include an MCP Security Gateway (tool poisoning detection, drift monitoring, typosquatting, hidden instruction scanning), Shadow AI Discovery (find unregistered agents), a Governance Dashboard, and a 12-vector PromptDefense Evaluator.

The policy format is YAML with Python-expression conditions:

rules:
  - name: block-destructive
    condition: "action.type in ['drop', 'delete', 'truncate']"
    action: deny

The verdicts are allow, warn, deny, escalate, or transform. The runtime is stateless, deterministic, and fails closed. The audit trail is a Merkle hash chain — each entry contains the SHA-256 of the previous entry.

Where AGT is stronger than PrimeSwarm

We do not pretend otherwise. AGT is broader.

Sandboxing and execution isolation

AGT's Agent Runtime provides 4 execution rings (Ring 0-3) with privilege isolation, session isolation with filesystem-level boundaries, a kill switch with step handoff and compensation, and saga orchestration for multi-step transactions with automatic rollback. PrimeSwarm does not provide container isolation. We said this on our own threat model page — we govern decisions, not processes.

Multi-agent identity and trust

AGT's Agent Mesh provides Ed25519-backed identity, DID-style agent credentials, trust scoring with decay and revocation, challenge-response handshakes, and delegated trust chains. PrimeSwarm governs a single agent's decisions. We do not have a multi-agent mesh.

SRE and reliability engineering

AGT's Agent SRE provides circuit breakers, error budgets, SLO-driven enforcement, chaos testing, and cascade detection across agent fleets. PrimeSwarm does not have this.

Framework support

AGT integrates with 14 frameworks: Microsoft Agent Framework, Semantic Kernel, AutoGen, LangGraph, LangChain, CrewAI, OpenAI Agents SDK, Claude Code, Google ADK, LlamaIndex, Haystack, Mastra, Dify, and Azure AI Foundry. PrimeSwarm has limited framework integrations.

Language SDKs

AGT ships SDKs in 5 languages: Python, TypeScript, .NET, Rust, and Go. PrimeSwarm has Rust and TypeScript.

MCP security

AGT's MCP Security Gateway detects tool poisoning, drift, typosquatting, and hidden instructions in MCP servers. PrimeSwarm has an MCP tool protocol but no security gateway.

Compliance framework mapping

AGT maps to OWASP Agentic Top 10, NIST AI RMF, EU AI Act, SOC 2, AARM Extended, and the Agentic Trust Framework. PrimeSwarm has the DGV test card suite but no formal mapping to these frameworks.

Scale and backing

AGT is backed by Microsoft, has 992 conformance tests, 29 ADRs, an OpenSSF Scorecard, CodeQL, Gitleaks, ClusterFuzzLite, and Dependabot. PrimeSwarm is an independent research lab with 89 DGV test cards.

Where PrimeSwarm is stronger than AGT

This is where it matters. AGT's own documentation — specifically LIMITATIONS.md and the OWASP mapping — explicitly states gaps that PrimeSwarm addresses.

Knowledge governance (AGT Limitation #7)

AGT says, in their own words:

"AGT governs agent actions. It does not govern the knowledge agents consume — the documents, databases, embeddings, and context retrieved during reasoning. AGT does not verify the provenance, freshness, or authorization of documents retrieved via RAG. AGT does not track which knowledge sources influenced an agent's decision. AGT does not enforce data classification labels on retrieved context."

PrimeSwarm has provenance enforcement. No legal, clinical, or financial assertion without a source receipt. The Continuity Ledger records which knowledge sources influenced each decision. For regulated industries — where an agent citing the wrong regulation or a stale clinical guideline is a liability — this is the gap that matters.

Memory governance (AGT ASI06 — Partial)

AGT's OWASP mapping marks ASI06 (Memory and Context Poisoning) as Partial:

"No dedicated memory-sandbox or context-integrity module. The audit hash-chain provides tamper detection for any persisted state. However, AGT does not yet sandbox agent memory stores."

PrimeSwarm has the Continuity Ledger — durable governed memory with Ebbinghaus decay tuned to domain relevance. Memory is governed, not just audited. Clinical memory decays differently from financial memory. Recent lab values matter more than six-month-old notes. The Ledger enforces this.

Fairness and bias

AGT does not mention fairness gates, Adverse Impact Ratio, the Four-Fifths Rule, or protected-class analysis anywhere in their specifications, OWASP mapping, or limitations document. PrimeSwarm has fairness gates with automatic Adverse Impact Ratio calculation and Pre-Effect Refusal when bias is mathematically detected. For a lending platform, a hiring tool, or a clinical triage system, this is not optional. It is the law.

PII sanitization

AGT mentions "prompt/content sanitization" and "blocked_patterns to catch PII in outputs." PrimeSwarm has a dedicated PII sanitization layer that strips SSNs, account numbers, and NPI from context before the agent processes them, replacing them with cryptographic hashes. The agent never sees the raw PII. AGT catches PII on the way out. PrimeSwarm removes it on the way in.

ONLY Lang vs YAML conditions

AGT's policy format is YAML with Python-expression conditions. The condition "action.type in ['drop', 'delete', 'truncate']" is code. A compliance officer cannot read it. A developer must write it.

PrimeSwarm's ONLY Lang is a purpose-built governance DSL with mathematical semantics:

harmony([0.5, 0.3, 0.2])
assert_bounds(0, 50, 0)
budget_limit(500)
escalate

A compliance officer can read this. The policy is the policy — not code that implements the policy. For an auditor who needs to verify that the governance script matches the regulatory requirement, ONLY Lang is a different category of artifact.

Cryptographic receipts vs Merkle hash chain

AGT has a Merkle hash chain audit log. Each entry contains the SHA-256 of the previous entry, making tampering detectable. This is good. It answers: "what happened."

PrimeSwarm produces cryptographic receipts — SHA-256 of script + input hash + gate state + residual + timestamp, forwarded to SIEM, replayable with the open-source verifier. This answers: "replay this exact decision and verify it yourself."

The difference: an auditor reading an AGT audit log trusts that the log is accurate. An auditor reading a PrimeSwarm receipt does not need to trust the receipt — they can replay the evaluation with the open-source verifier and confirm the result independently. For an auditor who needs to verify a decision without trusting the platform that made it, receipts are stronger.

Regulated-industry deployment model

AGT is a toolkit. You pip install it and integrate it yourself. The documentation is excellent. The examples are thorough. But the deployment is your problem.

PrimeSwarm is a deployed service with a 4-phase engagement model:

  1. Discovery — we map your existing decision flow to the governance pipeline
  2. Build — we define the ONLY Lang scripts that encode your compliance rules
  3. Deploy — we deploy to your VPC, on-prem, or private cloud
  4. Maintain — quarterly governance reviews, DGV re-verification, ONLY Lang script updates for regulatory changes

For a regulated industry that needs a deployed service with ongoing governance maintenance — not a toolkit — this is the difference.

Where they are comparable

DimensionAGTPrimeSwarm
Deterministic policyACS is stateless, deterministic, fail-closedGate is deterministic
Fail-closedYes — any error yields denyYes
Audit trailMerkle hash chainSHA-256 receipts + SIEM
HITL approvalApproval workflows with named approversHITL escalation with cryptographic sign-off
Open sourceMITMIT/Apache (verifier, compiler, cards)
Prompt injectionInput rails, 12-vector PromptDefense EvaluatorTPNN spatial constraints

The honest bottom line

AGT is the broader toolkit. If you want a general-purpose agent governance toolkit that integrates with 14 frameworks, runs in 5 languages, covers sandboxing, multi-agent identity, SRE, and 6 compliance frameworks — AGT is the stronger choice. We do not compete on breadth.

PrimeSwarm is deeper on the dimensions that matter to regulated industries. We fill gaps that AGT explicitly documents:

  • Knowledge provenance (AGT Limitation #7)
  • Memory governance (AGT ASI06 — Partial)
  • Fairness and bias (not in AGT)
  • PII sanitization before processing (AGT does it after)
  • Cryptographic receipts vs hash-chain logs
  • ONLY Lang as a compliance-readable DSL vs YAML conditions
  • A deployed service vs a toolkit

They could be complementary. AGT could handle the infrastructure layer — sandboxing, identity, SRE, framework integration — while PrimeSwarm handles the decision layer — fairness, provenance, PII, receipts, memory governance. Both are deterministic. Both fail closed. Both produce audit trails. Running PrimeSwarm's ONLY Lang policies inside AGT's policy engine is architecturally feasible.

What we are not claiming

  • We are not claiming PrimeSwarm is "better" than AGT. AGT is broader. We are narrower and deeper on specific dimensions.
  • We are not claiming AGT is insufficient. AGT is excellent at what it does. Its limitations document is more honest than most vendors' marketing.
  • We are not claiming ONLY Lang is universally superior to YAML conditions. For a developer, YAML is faster. For a compliance officer, ONLY Lang is readable. Different audiences.
  • We are not claiming receipts are universally superior to hash chains. For internal auditing, hash chains are sufficient. For external auditors who need independent verification, receipts are stronger.
  • We are not claiming the DGV test cards are superior to AGT's 992 conformance tests. They serve different audiences. AGT's tests verify the toolkit works. DGV cards verify the governance works.

What we are claiming

PrimeSwarm addresses gaps that AGT explicitly documents. For regulated industries — finance, healthcare, legal, accounting — those gaps are not optional. A lending platform without fairness gates is a lawsuit. A clinical agent without provenance enforcement is a malpractice risk. A legal agent without memory governance is a privilege violation.

If you are evaluating both, the question is not which is broader. The question is which gaps matter to your threat model.

Try it yourself

  • Run the governance gate in your browser at /try — write an ONLY Lang script, set the field values, press Run, see the receipt.
  • Browse all 89 DGV test cards at /dgv/registry — each one is a verifiable governance test.
  • See the full threat model at /threat-model — what we cover, what we don't, what we recommend for the gaps.
  • Read the AGT repository at github.com/microsoft/agent-governance-toolkit — it is excellent and we recommend you evaluate it honestly.

For the full verifier bundle — run all 89 cards locally and produce signed receipts — go to /try/bundle.


Back to Publications

Published by Only Institute