Question Methodology: How Our AFOQT Questions Are Written
How our AFOQT practice questions are written, generated, reviewed and retired — and exactly what our practice scores can and cannot tell you.
Last reviewed August 4, 2026
Every question here is original
We write our questions from the publicly described subtests. We do not have access to operational AFOQT content, we do not want it, and we would not publish it if it were offered to us. That is a legal and testing-integrity position, but it is also a practical one: operational forms change, and an item somebody half-remembered after sitting the test is as likely to teach you a wrong answer as a right one.
The rules we hold ourselves to:
- Create only original questions written from the publicly described skills and subjects.
- Never copy, paraphrase, scrape or reconstruct secure or recalled AFOQT items.
- Never source items from Quizlet, Reddit recollections, Amazon previews, commercial books or competitor banks.
- Never solicit remembered questions from users.
- Store provenance, author, technical reviewer, editorial reviewer, version and correction history for every item.
- No AI-generated item may publish without both technical and editorial approval.
If you have material you believe is recalled or copied from the real exam, please do not send it to us. If you think something on this site resembles a real item too closely, tell us through the corrections route and we will withdraw it while we investigate.
How an item is built
We write questions backwards, starting from the mistake the item is meant to catch. Before the prompt exists, the writer names the specific error a candidate makes on that skill — reversing the direction of an analogy, answering with the quantity the question did not ask for, taking a second percentage against the original total rather than the remainder, miscounting a row in a lookup grid.
That named error becomes one of the wrong answers. The remaining distractors each correspond to another identifiable wrong turn: a sound-alike word, an unsimplified form, a reciprocal, an adjacent cell. A distractor that nobody would plausibly choose is not a distractor — it wastes a slot and makes the item easier than the real thing, which is a quiet way of flattering the candidate rather than preparing them.
Every item then needs a short explanation giving the immediate reason, a full worked solution, and a rationale for each incorrect option explaining why it is attractive. The schema rejects an explanation that merely restates the answer, because a key without reasoning teaches nothing about the next question.
Visual items are generated, not drawn
Three subtests are visual, and for all three we generate the figure from data and compute the answer from that same data. This removes an entire class of defect: a practice bank cannot develop a key that no longer matches its picture, because there is only one source.
Table Reading
A grid is produced from a numeric seed, so the same seed always yields the same table. The answer is read directly out of the array the candidate sees. The four wrong answers are the cells immediately above, below, left and right of the correct one — precisely where you land after miscounting a single row or column, which is how this subtest actually punishes you.
Block Counting
A pile is described as a height map: a grid of column heights. That representation cannot describe a floating block, so every generated structure is physically valid. The number of blocks touching the marked block is computed by checking the four side neighbours at the same level plus the blocks directly above and below. Diagonal contact does not count, and neither does an edge or a corner.
Instrument Comprehension
An item is defined by three numbers: bank, pitch and heading. The artificial horizon and compass are rendered from those values, and the answer choices are generated by reversing one element at a time — bank direction, pitch sense, or the reciprocal heading. Those three are the mistakes candidates genuinely make, so those are the wrong answers worth offering.
Because each figure is drawn from numbers rather than shipped as an image, the same numbers produce an accurate text description for anyone using a screen reader. The description cannot drift from the picture, because neither is authored by hand.
The publication gate
Nothing is published without two named approvals
An item is served only when its validation status is approved and both a technical reviewer id and an editorial reviewer id are recorded against it. This is a function in the codebase that every question passes through, not a checklist somebody is trusted to follow.
We do not currently have named subject-matter reviewers, and we will not invent them to unlock the gate. The consequence is not that our exams run short — it is that they do not run at all. A form is refused rather than shortened, because serving eleven questions where the exam has twenty-five would misrepresent the exam and produce a diagnostic nobody should act on. Every launcher reports how many approved questions exist, which is currently none.
An item moves along this path, and the workflow in the code refuses to skip a stage or to leave a review stage without the corresponding reviewer recorded:
- Draft — written, with its target error, explanation and rationales in place.
- Automated validation — machine checks first, so a reviewer never spends attention on a defect a computer could have found: key matches an option, every distractor carries a rationale, the hint ladder is complete, and for figure items the answer is recomputed from the data that draws the picture.
- Technical review — a subject-matter reviewer confirms the answer is correct, the reasoning is sound and the distractors are genuinely wrong.
- Editorial review — a reviewer checks clarity, difficulty labelling, tagging and tone.
- Accessibility review — for any item with a figure, a reviewer confirms the text alternative is enough to answer the item without seeing the picture.
- Approved, then published — two decisions rather than one, each with its own audit entry. Approval says the item is sound; publication says it should be in rotation now. Both reviewer ids are recorded permanently against the item.
- Retired or suspended — withdrawn, with the reason recorded. A disputed key is suspended immediately and stays out of rotation until it is resolved.
Automated checks run before any of that: the key must match a listed choice, choice ids and text must be unique, an explanation of real substance is required, tags must be present, and computed items must agree with their own geometry. A malformed item fails the build rather than reaching a candidate.
What we record about every item
Each question carries its author, both reviewers, its validation status, an immutable version history, the sources behind any factual content, its copyright provenance, and a record of whether AI assisted in drafting it. After launch it also accumulates real statistics — how many candidates attempted it, what proportion answered correctly, the median time spent, and how often it was reported as unclear.
Those statistics are stored as null rather than zero until there is genuine data. An item that has never been attempted does not have a zero difficulty; it has no measurement, and displaying one would be inventing a number.
How we use AI, and where we do not
AI tools assist with drafting explanatory prose, suggesting candidate distractors, and writing the code that runs this site. We disclose that here rather than hiding it, because a study site quietly publishing unreviewed machine output is exactly what a candidate should be wary of.
What AI does not do is decide anything consequential. No answer key is accepted on a model’s word — for the visual subtests the key is computed from the geometry, and elsewhere a person verifies it. No factual claim about the exam is published without a registered source and a verification date. No author biography or credential is generated. No citation is accepted that has not been checked to exist.
The reason for that caution is specific to this subject. Language models are fluent and confident about details they have partly invented, including the number of questions on a subtest or the minimum score for a career field. On a study site those errors do real damage: a candidate revises the wrong material, or makes a decision about when to test based on a figure nobody checked. Our full position is on the AI content policy page.
What our score report means
After a test you get accuracy, a pacing breakdown, an unanswered count and performance by topic, with your weakest areas ranked. That is genuinely useful for deciding what to study next, and it is what the report is for.
This is your performance on our original practice material. It is not an AFOQT score, percentile or composite. The Air Force does not publish the conversion tables that would make such a calculation possible, so nobody outside the Air Force can produce one.
We use the words practice performance and diagnostic estimate deliberately, and we never describe output here as an AFOQT score, percentile or composite. Where a figure rests on a small number of questions we say so, because a percentage drawn from six items moves twenty points on a single lucky guess.
If we ever publish a percentile estimate, it will be because an independent validation study supports it, and it will carry its limitations prominently. Until then, treat the numbers here as a way to track yourself over several weeks rather than a prediction of your result.
Reporting an error
Every question carries a stable id, and every question can be reported using it. A report that names the item and explains why the key looks wrong is the single most valuable thing you can send us. Confirmed errors are corrected, the item is versioned, and the change is logged publicly on the corrections page — we do not edit mistakes away quietly.
Our sourcing and review standards are set out in the editorial policy, and who does the reviewing is on the authors page.
Who wrote this, how it was made, and why
- Who
- Written by the AFOQTPracticeTest.net editorial team. We do not yet have named subject-matter reviewers with verifiable credentials attached to this site, and we will not invent them — so this page carries no expert byline. See authors and reviewers for how review works today and what we are changing.
- How
- This page documents the actual production pipeline implemented in the codebase: the question data model, the deterministic visual generators, the validation schema and the publication gate. Nothing described here is aspirational — each control named below exists as code and is covered by tests.
- Why
- A candidate deciding whether to trust practice material needs to know where the questions came from, who checked them, and what the score means. Most test-prep sites do not answer those questions. This page does.
- Last reviewed
- August 4, 2026. Read our editorial policy or report an error.