Deep, practical, and verified — written for people who build with AI.
GSA's OneGov deals for ChatGPT, Claude, and Gemini end Sept 30, 2026, covering ~3.4M federal workers. What is known, what isn't, and how to prepare.
Why a reproducible result is the load-bearing requirement in federal AI work: contractor transition, audit access, ATO evidence, and defending a reported number a year later. Plus
An honest sort of what an AI engineer already knows that counts in federal work, what has to be relearned (accreditation, air-gapped operation, documentation, provenance), and what
A step-by-step guide for engineers and small firms entering federal work: UEI and SAM.gov registration, CAGE, NAICS selection, where opportunities are posted, and a ninety-day sequ
Unlimited, government purpose, limited and restricted rights explained for first-time federal vendors: technical data vs software, private expense, assertions, markings, and SBIR d
A buyer's guide to interrogating an AI accuracy number: what dataset, what split, how many runs, what spread, what baseline, does the metric match the decision, and can it be repro
A practical reading order for a uniform-contract-format RFP: Section M, Section L, Section C, Section B, Section I, and the format traps that get proposals rejected before evaluati
Warrant authority and its limits, how a CO differs from a COR and a program manager, who may direct contract work, constructive change, and how vendors should communicate with each
What CUI is under 32 CFR Part 2002, how it differs from classified, the marking rules, where it may be stored, what DFARS 252.204-7012 requires, and the 72-hour incident clock.
The simplified acquisition threshold is $350,000 and the micro-purchase threshold is $15,000 as of October 1, 2025. What each lane requires, how quotes work, and why this is the ea
How the Revolutionary FAR Overhaul works: class deviations, the new Part 40 security consolidation, the Part 39 rewrite, and the 52.2 to 52.4 clause renumbering.
A demo proves a model is learnable. Production needs evaluation, drift monitoring, retraining as a release, and an on-call owner. That is where pilots die.
A fluent, cited, wrong RAG answer is a symptom of one of six upstream layers. How to find which one is failing — the tests, in the order that isolates each.
The quote covers inference. The bill covers data preparation, evaluation, integration, and operation. A plain accounting of where AI project money goes.
When an off-the-shelf AI product wins, when it cannot, and the four questions that decide it: data, differentiation, integration depth, and ownership.
Regulated data forces four decisions early: residency, retention, access control, audit trail. What each one changes, and what breaks when you defer it.
Where an AI component lands in the Trust Services Criteria, the evidence a SOC 2 auditor will ask for, and what to instrument before the window opens.
Seams, strangler figs, characterization tests, and contract tests — how to replace a legacy system incrementally, and why big-bang rewrites keep failing.
Schema drift, silent nulls, late-arriving data, and non-idempotent tasks explain most weekly pipeline breakage — and each has a documented fix.
Filtering, hybrid search, index type, memory profile, operations, and lock-in — the six dimensions that decide whether a vector database choice holds up.
Uptime is green and the answers are wrong. What to log on every model call, how to sample for human review, and how to catch quality regression without labels.
Most anomaly detectors end up muted, and the cause is arithmetic. Why false-positive rate dominates, and how to design for the person carrying the pager.
Gaps, regime changes, and promotions break demand forecasts. Why the naive baseline is the experiment, and how to tell a real gain from a reported one.
Extraction versus generation, layout and table handling, and the span contract that keeps every extracted field traceable to a character range in the source.
Batch versus streaming, compared honestly: what exactly-once actually guarantees, what it costs in latency, and when a nightly batch is the right answer.
Lift-and-shift traps, egress and NAT charges, right-sizing data you are not collecting, and the licensing review that prevents a cloud migration cost overrun.
Multi-step tool-calling systems, where autonomy actually pays, the failure modes that kill agent projects, and why a deterministic workflow often wins.
The three CMMC levels, what a small defense contractor must actually do, and where the effort concentrates: scoring, POA&M limits, scope, and evidence.
The artifacts, the control inheritance, the shape of the timeline, and the early engineering decisions that decide whether an ATO takes months or years.
A compact language model post-trained to work through a problem step by step: what the term means, typical parameter ranges, and where SRMs are used.
What data belongs at DoD impact levels IL2, IL4, IL5 and IL6, why IL1 and IL3 do not exist, and what each level fixes — from the DoD Cloud Computing SRG.
An ATO is a named government official's decision to accept risk and let a system run. Who signs it, the SSP/SAR/POA&M artifacts and how inheritance works.
What an air gap actually means, why it differs from IL5 and IL6, why most AI tooling assumes a network, and a checklist for running a model without one.
RAG writes new prose from retrieved documents. Extraction returns a pointer into one document. How to tell which one your problem actually needs.
OSINT is intelligence produced from information anyone can lawfully obtain: the official definitions, the source categories, and the sourcing discipline.
What provenance means for an AI output, the difference between a real citation and a plausible-looking one, and how to tell which kind you have.
What precision and recall actually measure, why either alone misleads, what a false-extraction rate adds, and how to read a vendor benchmark critically.
Headless means running a model as a component with no user interface. Why buyers ask for it, and the seven machine-facing interfaces it still needs.
A hallucination rate is the output of a measurement procedure, not a property of a model. How to define it, what to count, who judges, and how wide it is.
GSA's OneGov deals put ChatGPT, Claude, and Gemini in front of federal agencies for around a dollar a year. What the deals cover, what they don't, and the real work that follows.
OMB Memo M-26-04 requires agencies to buy LLMs that meet two "unbiased AI" principles and to collect specific vendor documentation. A plain-language guide to the mechanics.
Claude Code, Codex, Cursor, Grok Build — by team.
Why every tool speaks Model Context Protocol.
Budgets, retrieval, caching, memory.
Mid-2026 pricing for the major tools.
A framework with worked examples.
When running local actually wins.
Cache mechanics and cost math.
Golden sets, LLM-judge, regression suites.
The June EO, M-25-21, GAAIA.
TSMC, Korea, and what it means for builders.