Methodology

The mechanism, the standards, and the research.

How the two lines are separated, where we actually stand against the independence parameters, what the published record says — and what a person still decides.

The premise

An agent that is confident and wrong is worse than no agent at all. The method exists to make confidently wrong the hardest outcome to produce.

CT Command runs program and portfolio work with agents. That is only defensible if the output can be traced, checked, and refused. So the platform is built around four commitments that apply to every deliverable it produces, and a human gate on anything that changes a client's system of record.

What follows is written with enough specificity that you can hold us to it.

The mechanism

One system. Two lines. A check that does not share the builder's assumptions.

Input

Client source documents

Contracts, plans, budgets, filings, policies and records the client actually provided.

Delivery line

Agents draft against a governed doctrine registry. Every claim carries its source.

Produces the work
Verification line

A different model from a different lab, with its own evaluation tooling. It forms its own reading of the sources and flags claims they do not support.

Judges the work
Supported

Human sign-off, then written back into the client's own tools.

Not supported

Abstain or escalate. Never fabricated.

Audit trail captured across both lanes, on the same pass

Assessed against IEEE 1012-2024, Annex C

Independent verification has three parameters. Here is where we actually stand.

The international standard for verification and validation defines independence along three axes. Most vendors claim the word without naming the test.

Technical independence
Met

The verification line uses a different model from a different lab and its own evaluation tooling. It never sees how the delivery line reasoned, only what it produced and what the sources say.

Managerial independence
Partial

The verification line selects its own findings and reports them without the delivery line's approval. It is not yet a separate organization. Full organizational separation arrives with scale.

Financial independence
Not yet

One company, one budget. Third-party attestation is the path, and we will say so plainly until it is in place.

By that standard we are internal IV&V today with strong technical independence, moving toward full independence as we grow. We would rather publish the assessment than the adjective.

The published record

The design is not a guess. It follows the published record on whether one model can check another.

The technique originates with Du, Li, Torralba, Tenenbaum and Mordatch, Improving Factuality and Reasoning in Language Models through Multiagent Debate, ICML 2024 (PMLR 235:11733–11763), which found that having three instances of a model critique each other over two rounds improved accuracy over a single instance across six benchmarks — for example 66.0% to 73.8% on factual biographies and 63.9% to 71.1% on MMLU. Those runs used three instances of one model, from one lab, on roughly 100 items per benchmark.

That matters, because the follow-up literature is not flattering. Zhang, Cui, Chen, Wang, Zhang, Wang, Wu and Hu, Stop Overvaluing Multi-Agent Debate (arXiv:2502.08788), evaluated five debate methods across nine benchmarks and four foundation models and found they often fail to beat a single agent using simple chain-of-thought or self-consistency, even while burning far more compute.

Their finding is what our architecture is built on: model heterogeneity was "a universal antidote to consistently improve current MAD frameworks." Agents drawn from different models, not copies of one, are what makes the check real.

So we do not run a model against itself and call it verification. The verification line is a different model from a different lab, it never sees the delivery line's reasoning, and where the sources do not support a claim the answer is an abstention, not a guess.

This is external research on verification generally. It is not a benchmark of CT Command.

The Four Commitments

What every deliverable has to clear.

01

Sourced

Every deliverable is built from named source documents.

Nothing is produced from general recall. Work is grounded in the documents a client actually gave us — contracts, plans, budgets, filings, policies, meeting records. Each load-bearing claim traces back to the document and the section it came from, so a reader can follow any statement to its origin rather than taking it on faith. A claim without a source is treated as a defect, not a stylistic choice.

Inputs
Client source documents · scope of the request
Outputs
Draft deliverable · claim-level source trail
02

Independently verified

A second agent line, running a different model from a different lab, reviews the output against the same sources.

The reviewer is walled off from the builder. It does not see the delivery line's reasoning, only the finished output and the source set, and its job is to find claims the sources do not support. This is the independent verification and validation model used in safety-critical software: the value of the check comes from the separation, not from the checker being smarter. Disagreements surface as flagged claims, not silent edits.

Inputs
Draft deliverable · source set
Outputs
Verification result · flagged claims
03

Abstain rather than fabricate

A faithfulness floor governs both lines.

When an answer is not supported by the source documents, agents say so, or escalate to a person. They do not fill the gap with something plausible. In practice this means a deliverable can come back with an open question attached to it, and that is the intended behavior — a marked gap is information a manager can act on, while an invented answer is a liability that surfaces later at the worst possible time.

Inputs
Question · available sources
Outputs
Supported answer, marked gap, or escalation
04

Attested and isolated

No agent reaches a client without a signed security attestation. Each client runs in a firewalled instance.

No agent reaches a client until it has been red-teamed against the OWASP, NIST and MITRE frameworks and carries a signed attestation describing what it was tested against. Client environments are isolated on the read side; server-side tenant binding for write capture is on the critical path, and no client instance ships until it closes. The attestation travels with the agent, so the security posture of a given piece of work is inspectable rather than asserted.

Inputs
Agent release candidate · adversarial test pass
Outputs
Signed attestation · firewalled client instance

Human sign-off

What a human still decides.

Automation earns its place by removing the work no one should be doing by hand. It does not earn the authority to commit on a client's behalf.

Human gate

Agents draft and verify; a person commits

The two agent lines produce and check the work. They do not decide that the work is finished. A named person reviews the result and takes responsibility for it before it leaves the platform.

Human gate

Every write passes a human gate

Nothing an agent produces writes back into a client's system of record on its own. Changes are staged, reviewed, and released by a person, so the client's source of truth only ever moves when someone chose to move it.

Human gate

Consequential actions escalate, not execute

Actions that are irreversible, externally visible, or materially consequential are routed to a person rather than performed. The platform's default in an ambiguous situation is to stop and ask.

See the method applied to work you actually have.

Bring a real program, a real portfolio, or a real question — we'll show you what the platform produces and how it holds up.