AI Modernization · Federal & Commercial
Results Based AI Modernization.
We’ve done the frontier work in workflow automation, agentic teams, governance and traceability, model economics, and design. We help you move to better AI at the pace your culture can carry and the speed your market demands.
The incumbent model
What AI consulting can feel like.
-
A large bill, up front.
Six- and seven-figure engagements priced before a single workflow works. You fund the transformation, then hope for it.
-
Top-down, by “experts”.
A strategy handed down from people who have never done your job, mapped onto a workflow they’ve only seen in a slide.
-
Your problems, their education.
Your hardest cases become the vendor’s training ground: you pay a premium to teach them the work you already understand.
It shows in the industry’s own numbers.
We built CoA to be the opposite.
Our model
Three ways we make AI real for you.
Each is detailed in full below.
- 01
AI Transformation
Step by step, with people who have actually built and shipped AI — across technology, workforce, culture, and cost — matrixed to your industry, your compliance regime, and your competitive reality. We prove value on a real workflow before you fund the rollout, so the plan is grounded in your data, not a slide.
See AI Transformation in full - 02
Model Economics
We find where AI earns its place in your organization, then route each task to the smallest model that does it well, so run cost stays low without giving up quality. Task categorization and sample testing set the mix; measured accept/reject rates keep it honest.
See Model Economics in full - 03
Traceability & Accountability
The metrics that tell you whether AI is actually helping: hallucination rate, prompt efficiency, and helpfulness measured as accept, edit, and reject rates on assistive output — every action on the record, mapped to the frameworks an evaluator expects. The same instrumentation we run on our own products.
See Traceability & Accountability in full
01AI Transformation
Bridge where you are to where you want to be.
Not a tool install — a workflow redesign across four axes: technology, workforce, culture, and cost. Every component below is work we have delivered.
Technology
The systems that make AI real.
-
Prove-before-you-fund pilot
A real, working AI product on your actual workflow and data in weeks — value shown with live metrics, not a deck, before you fund the rollout.
-
Domain-grounded model build
The assistant is built on your own corpus — case law, benefits rules, clinical data — so it is accurate and cheaper per accepted output than a generic chatbot.
-
Identity & access foundation
The ICAM / SSO substrate that lets AI reach authenticated systems safely — grounded in our IRS and VA identity work.
-
Production delivery & integration
The proven pilot becomes a real, deployed, authenticated application wired into your existing systems, not a demo left on a laptop.
Workforce
The people who run it.
-
Human-in-the-loop workflow redesign
We re-engineer the decision so AI augments the person — with review, override, and escalation — drawn from how Adjudicate handles case management and decision writing.
-
Enablement from operator practice
We teach your staff to run the AI-augmented workflow — review output, catch slop, escalate — from how we operate every day, not from a generic curriculum.
Culture
The trust that makes it stick.
-
AI readiness & opportunity map
A diagnostic that inventories where AI creates real return — and where it should not be used — scored against your data condition, compliance regime, and staff capacity.
-
Anti-slop output-quality gate
An explicit acceptance bar every AI output clears before it reaches a decision or a citizen. It turns skeptical staff into adopters and satisfies an auditor at the same time.
Cost
The economics that keep it sustainable.
-
Vendor-neutral tooling & cost instrumentation
Models and tools chosen independent of any single vendor and matched to each workload, with measured token efficiency. Run cost is shown as cost-per-accepted-output, not asserted.
-
Governance & traceability instrumentation
The measurement layer — hallucination rate, accept/reject rates, prompt efficiency, per-decision audit trail — stood up from day one, the same one we run on our own products.
02Model Economics
The right model on the right task.
Most AI bills are one expensive model doing every job. We categorize your tasks, sample-test each tier, and route work to the smallest model that does it well, so cost tracks the work, not the ceiling.
Haiku 4.5
Classify, extract, route, tag
$1 in $5 out / 1M tokens
Sonnet 5
Summarize, draft, most work
$3 in $15 out / 1M tokens
Opus 4.8
The hard cases
$5 in $25 out / 1M tokens
Fable 5
The rare hardest problems
$10 in $50 out / 1M tokens
A theoretical estimator
Set the share of your tasks each tier handles. We show blended cost per task versus running everything on Opus. In our experience, task categorization and sample testing set the real mix for your organization.
than all-Opus
On a representative assistive task — ~10K input + 2K output tokens. Anthropic list pricing, 2026 snapshot. Illustrative; your real mix and savings come from sample-testing your own workloads.
03Traceability & Accountability
Metrics attuned to where AI is and is going.
The governance we build into every product — from the one rule underneath it to the frameworks an evaluator expects.
AI proposes. People decide.
Every output is a proposal a person accepts, edits, or rejects. The model never acts on its own. Traceability is how you prove it held. The three questions leaders actually ask:
- Hallucination rateThe rate of confidently-wrong output — measured, not assumed away.
- Prompt efficiencyOutput quality per prompt: rework, retries, cost and latency.
- HelpfulnessAccept, edit, and reject rates on assistive output.
Every way an LLM fails, and the control that contains it.
- Hallucination / fabricated factsOWASP LLM09 · NIST MEASURE
A contextual grounding check scores grounding and relevance against the source on a configurable 0–0.99 threshold and blocks answers the source does not support; every answer must carry citations. Grounding proves an answer is supported by the record, not that the conclusion is correct.
- Prompt injection / jailbreakOWASP LLM01
Prompt-attack detection runs on both input and output, the guardrails apply independently of which model is used, and tools are granted only least-privilege scopes.
- Sensitive-data / PII disclosureOWASP LLM02
PII is detected and masked in the model’s input and output, with data encrypted in transit and at rest and the case kept inside the deployment’s account and region. Masking only covers what the model sees, so logs and traces, which can retain raw text, are protected separately by data-protection policies and least-privilege access.
- Overreliance / automation biasOWASP LLM09 · NIST MANAGE
Each proposal leads with its evidence and a confidence so the reviewer engages with the basis; high-impact steps require explicit confirmation; and accept / edit / override rates are monitored so rubber-stamping is detectable.
- Excessive agencyOWASP LLM06
Least-privilege permissions, narrowly scoped tools, and explicit human approval before any outbound action.
- Bias / inconsistent outcomesNIST MEASURE (bias)
Supports disparate-impact testing across cohorts via cohort-stratified evaluation and golden sets, with fairness metrics and periodic equity audits configured per deployment; rule changes are tested before they apply.
- Model / data drift & degradationNIST MEASURE → MANAGE
LLM output quality and drift are tracked by the evaluation harness and reviewer acceptance; tabular and feature drift by Model Monitor; regression evaluations run on every change, with alarms on degradation.
- Supply chain / model provenanceOWASP LLM03
Models come from a vetted, managed catalog with known provenance, and dependencies and artifacts are reviewed.
- Improper output handlingOWASP LLM05
Every model output is treated as untrusted: generated content is validated and never auto-executed.
- Unbounded consumption / cost & DoSOWASP LLM10
Rate limits, per-tenant quotas, and throughput controls cap usage, and alarms fire on abnormal demand.
Measured per response, per change, and continuously.
- Golden evaluation setsPer release
Curated cases with known-correct answers; every model or prompt change is re-scored against them. Regressions surface before release, not in production.
WatchesCorrectness vs. known-good answers
- Contextual grounding + citation checkPer response
Every answer is scored against its retrieved source, and every citation is checked against the record. Unsupported output is flagged or blocked before anyone relies on it.
WatchesIs each output supported by its cited source?
- Scaled dry-run / diffPer change
A change runs across many cases in read-only mode and shows up as a before→after diff — you see exactly what it would do before it touches live work.
WatchesWhat a change would do, before it applies
- Human acceptance telemetryContinuous
Accept, edit, and reject rates, tracked per feature and per reviewer. Rising edits flag a model that needs attention; suspiciously few flag rubber-stamping.
WatchesAccept / edit / override rates (over- and under-trust)
- Drift & quality monitoringContinuous
Inputs and outputs are watched for drift from the validated baseline, and an alarm fires the moment a metric crosses its threshold.
WatchesDrift & quality degradation vs. baseline
- Red-team / adversarial testingPer release
Prompt injection, jailbreaks, and edge cases, thrown at the system on purpose — we find the failure modes before real inputs or bad actors do.
WatchesPrompt injection & failure modes
Mapped to the standards an evaluator expects.
- NIST AI Risk Management FrameworkAI RMF 1.0 · 2023
Our controls map to the framework’s four functions and its Generative AI Profile (NIST AI 600-1, 2024). NIST AI RMF is a voluntary framework; mapping to it shows alignment, not certification.
- Govern
- Map
- Measure
- Manage
- GenAI Profile (600-1)
- OWASP Top 10 for LLM Applications2025
Each guardrail in the matrix above maps to the specific LLM application risk it addresses, so you can see which control answers which weakness.
- LLM01 Prompt Injection
- LLM02 Sensitive-Info Disclosure
- LLM05 Improper Output Handling
- LLM06 Excessive Agency
- LLM09 Misinformation
- LLM10 Unbounded Consumption
- Federal AI policyOMB M-25-21 / M-25-22
Built to current federal guidance on agency AI use and acquisition, OMB M-25-21 and M-25-22 (both Apr 3, 2025), under EO 14179. Whether a deployment satisfies it is the agency CAIO’s determination.
- Chief AI Officer
- High-impact AI practices
- Human oversight & appeals
- Model & data portability
- VA Trustworthy AI Framework2023
Aligns with VA’s six trustworthy-AI principles. Claims processing is a named VA AI operational area, and VA has adopted OMB’s high-impact-AI definition.
- Purposeful
- Effective & Safe
- Secure & Private
- Fair & Equitable
- Transparent & Explainable
- Accountable & Monitored
Every high-impact-AI practice, and how we meet it.
OMB M-25-21 §4(b) and M-25-22 (Apr 2025), under EO 14179. Whether a deployment satisfies it is the agency CAIO’s determination.
- Pre-deployment testing Evaluation harness, golden sets, and a scaled dry-run before any change applies
- AI impact assessment Grounded, auditable proposals and outcomes by issue type provide the evidence a written impact assessment draws on
- Ongoing monitoring Drift and accuracy monitoring plus accept / edit / reject telemetry
- Human training & assessment Role-specific guidance for adjudicators on what the AI does, its limits, and how to override it
- Human oversight & accountability The AI proposes and a reviewer decides, accepting, editing, or rejecting every output and signing the work
- Remedies / appeals The veteran keeps the full statutory appeal path (higher-level review, the Board, CAVC); AI involvement is disclosed so the decision stays contestable
- End-user & public feedback Reviewer and end-user feedback is collected and folded into evaluation and threshold tuning
- Vendor-lock-in protection (M-25-22) Configurable by model and provider, self-hostable, and portable
Start where it pays off