Technical SEO Audits Past the Crawl Limit: Judging 1,200 Pages in Three Minutes
Crawling is solved and cheap. Judging what the crawl found has always been human hours, which is why audits get quoted in weeks and sampled instead of finished. Here is the crawl versus judge split,...
A technical SEO audit is two jobs, not one. Crawling is solved, fast and nearly free. Judging what the crawl found (is this page thin, does it duplicate that one, what should happen to it) has always been human hours, which is why audits get quoted at 20 to 200 hours.1 That second half is the part that just got cheap.
The short version
- A crawler tells you what is on a page. It cannot tell you whether the page is any good, whether it duplicates another, or what to do about it. Every hour in an audit quote lives there.
- AgencyAnalytics puts an in-depth audit by an experienced SEO at 20 to 200 hours.1 That is the honest anchor, and the reason audits get sampled rather than finished.
- BoringToolsKit published a run that crawled 1,204 pages and returned 4,816 typed judgments in under three minutes, at $0.0048 for the weekly triage run, on jev-1.13.0. The crawling was done by LibreCrawl, not by the model.2
- Seven jobs in a standard audit reduce to one written question asked repeatedly, in one of three shapes: a yes or no, a pick from a fixed list, or a position on a named scale.
- How you ask beats which model you pick. On identical held-out data, one judgment split into five questions scored 95.0 percent against 89.4 percent for the same judgment asked once, while a frontier model moved the opposite way.5
- The cheap model is less accurate, and you should know by how much. Across 31,500 independently scored decisions it hit a Decision Score of 67.8 on pick-one questions against 74.1 for Gemini 3.8 Flash, at roughly a twenty-eighth of the price.6
- The real comparison is not model against expert. It is judging every page imperfectly against judging 50 well and guessing about the rest.
The two halves of an audit, and only one of them was ever hard
Run a crawler against a site and you get status codes, titles, meta descriptions, canonicals, headings, word counts, response times, inlink counts, crawl depth and a redirect map. Every one is a fact read off the page. No judgment is involved, which is why software has done it well for over a decade.8
Now look at what is actually in the deliverable. Which of these 1,200 pages says nothing the other 1,199 do not. Which pairs compete for the same intent. Which titles would disappoint the person who typed the query. Which of the 40,000 candidate internal links is worth writing. Which old URL each new URL replaces. None of that is a fact on the page. Each is a comparison, a judgment, or both.
That split explains the shape of the discipline. Crawler licenses are cheap because crawling is a commodity. Audit quotes are expensive and slow because judgment is not. If you want the mechanics of a typed decision model before you go further, we wrote the plain-English version in the Jev model explained.
Why audits are quoted in weeks
Four published figures tell you where the time goes.
AgencyAnalytics states that an in-depth audit by an experienced professional or an agency takes 20 to 200 hours.1 That is a ten-fold range, and the spread is less about site complexity than about how much of the site anyone actually reads.
A guest piece on the Screaming Frog blog breaks an enterprise audit into per-task hours: roughly 10 for the technical audit, 5 to 10 for content, about 5 for links, 25 for the content roadmap.4 The crawl itself is a small slice of every one of those numbers.
Seologist puts redirect mapping and URL alignment at 10 to 25 hours inside a migration, and says a 500-page site can often be mapped, redirected and tested within days while 50,000 indexed URLs needs weeks of phased redirection.3 That is the only published hours band for redirect mapping anywhere, and it scales with page count because the judgment is per URL.
For labeling work, an independent practitioner measured manual search intent tagging at 30 keywords per hour.15 That is what careful per-item judgment costs.
Do the arithmetic
What does this job cost, three ways?
List prices only, as published in September 2026. No benchmark claims, no projected savings. Change the job and watch which engine the arithmetic favors.
- Typed decision model$0.1680
Jev list price. Returns a value and a probability, no text.
- Small frontier model$1.44
gpt-5-nano list price. The honest comparison, and it is close on input.
- Frontier model$84.00
A mid-tier frontier model at $3 in and $15 out, the usual default.
Two things usually surprise people. Against a small model the input price is close, so the gap comes from free output and from answering several questions in one pass. Against a frontier model the gap is large enough that it changes what you are willing to run across a whole site rather than a sample.
Set those against the ceilings that force sampling. Screaming Frog's free tier caps at 500 URLs.8 Semrush caps a single Site Audit campaign at 20,000 pages on its mid tiers, with unused monthly crawl budget expiring rather than rolling over.9 Sitebulb Lite caps at 10,000 URLs per audit.21 Lumar does not publish pricing at all.10 The Search Console API returns at most 25,000 rows per request and states plainly that it does not guarantee all rows, only top ones.16
So the data arrives truncated before anyone reads it. Then a person reads a sample of the truncation. That is the state of the art, and nobody selling crawl seats will say so.
What 1,204 pages in three minutes actually measured
BoringToolsKit published a technical audit run in September 2026 with the numbers attached. It covered 1,204 pages and produced 4,816 typed judgments in under three minutes, at $0.0048 for the weekly triage run, on model version jev-1.13.0. Total model cost for the full pipeline came in under a dollar. Crawling was done by LibreCrawl, self-hosted, and every deployed fix was re-checked in a real browser at desktop and mobile widths.2
Those are BoringToolsKit's own published measurements, not an independent replication. Read them as a documented run by the people who built the pipeline.
The division of labor is the part worth copying. The crawler still crawls. The model never fetches a URL, counts affected pages, pulls Search Console, estimates engineering effort or approves a fix.13 It only answers questions about text something else already fetched. And 4,816 over 1,204 is four judgments per page, so each page got roughly four narrow questions rather than one big one. Hold that thought.
The speed comes from batching, not raw throughput. Questions sent in a single call are evaluated in parallel, so adding questions per page costs tokens and almost no wall-clock time.7 That changes what you ask. You stop rationing questions and start asking every page everything.
Which audit jobs reduce to a typed question
This is the useful part. Seven jobs that currently eat audit hours have an exact question shape, and writing the question down is most of the work.
| Audit job | Question type | The exact question | What it returns |
|---|---|---|---|
| Thin or mass-produced page | Score, 1 to 10 | "Score this page 1 to 10 for saying something the other pages do not" | A position on a named scale, with probabilities |
| Title and intent fit | Yes or no | "Would someone searching this query expect this title?" | A probability between 0 and 1 |
| Query to page mapping | Pick one | "Which of these 10 URLs should rank for this query?" | One option, with probabilities across the set |
| Keep, update, merge or remove | Pick one | "What should happen to this URL?" | One label from Keep, Update, Rewrite, Merge, Redirect, Remove, Human review |
| Redirect mapping in a migration | Pick one | "Which new URL replaces this old one?" | One target URL, with confidence |
| Cannibalization | Yes or no | "Do these two URLs answer the same search intent?" | A probability per pair |
| Internal link candidacy | Yes or no | "Is there an honest reason to link page A to page B?" | A probability per ordered pair |
Three notes on that table.
Pick-one jobs need the option set defined up front. A Choice question accepts up to 255 options, plenty for a label set and nowhere near enough for "which of my 12,000 URLs."7 Query to page mapping works only after cheap retrieval has cut the candidates to ten or so. The retrieval is code. The judgment is the question.
Yes or no jobs are pairwise, so the count explodes. The published arithmetic for an internal link map is 586 pages by 15 candidates each, or 8,790 calls.14 That is only survivable because something already cut 15 candidates out of 585. We covered the prefilter in the internal linking audit that runs in 45 seconds, and the same pairwise logic in keyword cannibalization as a yes or no question.
Label sets are a design decision, not a default. The classic four-way intent split is not the only option, and Content Harmony's nine SERP-aligned types fit a lot of catalogs better.19 We took that apart at volume in classifying search intent for 50,000 keywords.
One more job belongs here, carefully framed. "Does the structured data describe what is visible on the page?" is a clean yes or no. Treat it strictly as a data-integrity check: your markup should not claim a price, a rating or an author the rendered page does not show. It is not a claim that structured data improves how often an AI system cites you.
How you ask beats which model you pick
This result should change how you build an audit, and it is not a vendor number.
An independent write-up ran 1,000 held-out emails through the same job two ways. Asked as a single question, the typed decision model scored 89.4 percent. Split into five questions and composed in code, the same model on the same data scored 95.0 percent. Claude Haiku ran the opposite way: 94.2 percent on its best single signal, 93.2 percent composed.5
That is a 5.6 point swing from decomposition alone, on identical data, with no model change. The direction is not universal, which is the interesting half. A frontier model asked one rich question reasons its way toward an answer internally. A typed decision model cannot, so you decompose outside the model and combine the pieces in your own code.
Apply that to an audit and the design falls out. Do not ask "is this page thin." Ask four narrower things: does it say anything its siblings do not, does it answer the query it targets, does it have a reason to exist separate from the template, would a reader get what the title promised. Combine those with rules you wrote and can defend. Four judgments per page is what the BoringToolsKit run averaged, and a reasonable starting shape.2
Note the boundary. That is a non-SEO domain, so do not carry the percentages into an SEO deck. Carry the lesson: question design is the lever, and it is the one nobody is selling you.
The accuracy trade, stated plainly
Do not use a cheap judgment layer thinking it matches a frontier model. It does not, and the gap is measured.
Jevals scored 31,500 decisions across seven models, 300 questions per task, each asked five times. It reports a Decision Score where 100 is perfect and 0 is guessing the base rate. On pick-one questions Jev returned 67.8 against Gemini 3.8 Flash at 74.1. On yes-or-no questions it was 69.0 against 73.0. Jev did this at about one twenty-eighth of the price.6
So the trade is roughly six accuracy points for a 28x cost reduction, and whether that is a good deal depends on how many times you make the decision and what happens when it is wrong. For a judgment you make 8,790 times, where the output is a queue a person reviews and nothing publishes itself, it is obviously good. For a one-off decision on a revenue page, with a redirect going live at midnight, it is obviously bad. The skill is knowing which of your audit jobs is which.
Three more limits belong in the same breath.
It does not explain itself. System One models return a decision and a probability, not a reason.11 You never get "this page is thin because the top third repeats the category description." You get a number, and if it needs defending to a client, a person writes the defense.
Zero hallucination is a schema property, not a correctness property. Type errors and out-of-schema values are impossible by construction. That is not the same as the answer being right, and the launch coverage says the zero figure is not empirical.12 Confidence summarizes the probability distribution. It does not guarantee your diagnosis.13
Confidence needs a policy or it is trivia. Three bands: act automatically above the high band, queue the middle for a person, drop the low band rather than averaging it in. One open-source SEO implementation flags anything under 0.55 for human review instead of folding it into a score.18 Pick your bands against a labeled sample of your own pages, never against a number you read in a blog post, including this one.
Sampling versus coverage is the real comparison
Every objection lands as "a model is not as good as an experienced SEO." True, and not the comparison being made. On a 1,200 page site, the current method reads 50 to 200 pages carefully, pattern-matches the rest, and writes a deck. The alternative reads all 1,200 imperfectly, then hands a person a ranked queue of the ones where the judgment was uncertain or the stakes were high.
An experienced auditor beats a cheap model on any single page. That auditor plus complete coverage beats that auditor plus a sample, because the pages that break a site are rarely the ones a sample surfaces. The 900-page tail of a catalog is where the near-duplicates live, and nobody has ever read it.
Complete coverage also makes your thresholds testable. Label 200 pages by hand, compare the model against your labels, and tune the bands until the disagreement rate is one you can live with. You cannot calibrate a sample against itself.
Coverage also changes what gets found. SEOParity reports that a 10 to 15 percent gap in redirect coverage during a migration causes measurable loss.20 A 10 percent gap is what sampling produces on a large migration. The fix is not a better auditor. It is reading the whole list.
What we still do by hand, and why
A cheap judgment layer does not automate the job. It moves the human up the stack, and four things stay firmly human.
The diagnosis. A model tells you 340 pages score low on distinctiveness. It will not tell you all 340 came from the same 2023 template migration, and that fixing the template fixes all of them. Root cause is a pattern across findings, and the model sees one page at a time.13
The prioritization call. Ranking fixes means knowing which pages make money, who owns the CMS, what ships next sprint and what the board cares about this quarter. None of that is in the crawl, and engineering effort in particular is not something the model estimates.13
The client-facing write-up. A decision model is structurally incapable of prose. The architecture that works is deterministic code, then typed judgment, then a frontier model or a person where generation is required, then a verification pass, then the CMS.17 Not one model instead of another. Each doing what it can actually do.
Every irreversible action. Deleting pages, publishing redirects, changing canonicals on revenue templates. Human approval below your confidence threshold, and on revenue pages regardless. There is no version of this where a probability pushes a 301 to production on its own.
Ninety second check
Is your job decision shaped?
Five questions. One no is enough to make this the wrong tool, which is worth finding out before you wire anything up.
Can you write down every possible answer before you run it?
Is the output a decision rather than something a person will read?
Does it happen often enough that doing it by hand hurts?
If a call is wrong, can you undo it cheaply?
Can you live without a written reason for each call?
Answer the five above and you get a straight verdict here.
Where this fits: run the cheap judgment across everything for triage, then send the failures and the uncertain band to a frontier model or a person.11 Broad and cheap for coverage. Deep and expensive for the cases that earned it.
That is the shape of the work. The crawl was never the deliverable and neither is the spreadsheet. Our technical SEO work is built on this split, and the wider stack is in the AI and LLM SEO guide.
Frequently asked questions
How long does an SEO audit take?
AgencyAnalytics puts an in-depth audit by an experienced SEO professional or agency at 20 to 200 hours. A guest breakdown on the Screaming Frog blog splits an enterprise audit into roughly 10 hours for technical, 5 to 10 for content, about 5 for links and 25 for the content roadmap. Most of that is per-page judgment, not crawling, which is why the range widens with page count.
What does a technical SEO audit cost?
AgencyAnalytics publishes bands of $500 to $1,000 for a basic audit under 20 pages, $2,000 to $5,000 intermediate, $5,000 to $10,000 advanced and $10,000 to $20,000 enterprise, with agency hourly rates of $50 to $150 and up. The price tracks the hours, and the hours track how many pages a person reads.
Can AI do a technical SEO audit?
It can do the judgment half at scale and not the rest. A typed decision model answers questions about pages a crawler already fetched. It does not crawl, count affected pages, pull Search Console data, estimate effort, explain its reasoning or approve fixes. Diagnosis, prioritization, the write-up and every irreversible action stay with a person.
How many pages can a crawler audit at once?
It depends on the license. Screaming Frog's free tier stops at 500 URLs. Sitebulb Lite caps at 10,000 URLs per audit. Semrush caps one Site Audit campaign at 20,000 pages on its mid tiers, and unused crawl budget does not roll over. Lumar does not publish pricing. The Search Console API returns at most 25,000 rows and does not guarantee all of them.
Is a cheap decision model as accurate as a frontier model?
No. In an independent benchmark of 31,500 scored decisions across seven models, Jev scored 67.8 on pick-one questions against 74.1 for Gemini 3.8 Flash, and 69.0 against 73.0 on yes-or-no questions, at roughly one twenty-eighth of the cost. That trade works behind a review queue for a judgment you repeat thousands of times, and not for a one-off call on a revenue page.
Does splitting an audit question into smaller questions help?
Usually, for a typed decision model, and by a lot. On 1,000 held-out emails, one judgment asked as a single question scored 89.4 percent while the same judgment split into five and composed in code scored 95.0 percent. A frontier model moved the opposite way on the same data: 94.2 percent single, 93.2 composed. Non-SEO domain, so take the direction as the lesson, not the numbers.
Should I let a model decide which pages to delete?
Not on its own. Use it to read and rank every page, then work the result in bands: act on the confident ones as a reviewed batch, put the uncertain middle in front of a person, drop the low-confidence tail. One open-source implementation flags anything under 0.55 confidence for human review. Deletion and redirects are irreversible, so a probability should never be the last step before production.
What should a site audit checklist actually contain?
Two lists, kept separate. The crawl checklist covers status codes, canonicals, indexability directives, hreflang, redirect chains, sitemap coverage, render parity and response times. Software answers all of it. The judgment checklist should be written as questions rather than checks, because each one is a comparison two people can disagree about.
- AgencyAnalytics, "How Much To Charge For An SEO Audit," 14 August 2025. agencyanalytics.com
- BoringToolsKit, "SEO Audit Cost 2026," updated 19 September 2026. boringtoolskit.com
- Seologist, "How Many Hours Should I Estimate For Site Migration SEO," published 14 March 2025, updated 4 December 2025. seologist.com
- Mark Howser on the Screaming Frog blog, "How To Do An Enterprise SEO Audit The Right Way." screamingfrog.co.uk
- XenoSpectrum, "Jev, TypeSafe and BERT classifier decomposition," 20 September 2026. xenospectrum.com
- Jevals, independent System One model benchmark, 31,500 scored decisions across seven models, 18 September 2026. jevals.com
- Valyu on dev.to, "How to use Jev: a practical guide to TypeSafe's System One model," September 2026. dev.to
- Screaming Frog, SEO Spider pricing and free tier limits. screamingfrog.co.uk
- Semrush Knowledge Base, "How many pages can I crawl in an audit." semrush.com
- Lumar, pricing page, fetched 21 September 2026. lumar.io
- Arize AI, "TypeSafe Jev as an LLM judge," September 2026. arize.com
- MarkTechPost, "TypeSafe AI releases Jev," 19 September 2026. marktechpost.com
- Screpy, "What a System One model does not do in SEO issue triage," 2026. screpy.com
- Ryze, "Jev for SEO: nine jobs that reduce to typed questions," 2026. get-ryze.ai
- Niko Alho, "Intent classification with AI," published 20 May 2026, updated 18 July 2026. nikoalho.fi
- Google Search Console API, Search Analytics query reference, rowLimit. developers.google.com
- mean.ceo blog, "System One models in a content pipeline," 2026. blog.mean.ceo
- jevseo, open-source SEO pipeline, README confidence gate. github.com
- Content Harmony, "Classifying Search Intent." contentharmony.com
- SEOParity, "Redirect Map Site Migration." seoparity.com
- Sitebulb, pricing and per-audit URL limits, fetched 21 September 2026. sitebulb.com
Figures checked on 21 September 2026. The benchmark numbers are days old and the field is moving weekly. Calibrate your thresholds against your own labeled pages before you trust anyone's percentage, including ours.
Audit stuck in a spreadsheet?
Send us your domain. We run the judgment calls across every URL, not a sample, and hand back a sorted fix list with the uncertain pages flagged for a human.
- Cannibalization, intent and thin-page calls across the whole crawl
- Redirect and internal link maps you can ship
- Scoped estimate within 48 hours
Want to discuss ai search optimization for your business?
Start a project and we'll talk through where you are, what's working, and the highest-leverage moves for the next 90 days.


