Back to Blog
Agency
9 min read

What 100 Client-Submitted AI Audits Taught Us About LLM SEO Tools

Matt Weitzman
Senior SEO Strategist & Co-Founder
What 100 Client-Submitted AI Audits Taught Us About LLM SEO Tools

Picture this: a new client hops on an intro call, already holding an "audit." Sixty-three slides. Color-coded severity scores. A section called "Critical Issues" with fourteen line items. They want to know why their old agency hasn't fixed any of it yet. You skim the document. Within five minutes you've found three "critical" errors that don't exist, a ranking claim with no source, and a schema recommendation that would actually break their structured data. That's not a one-off story. After reviewing more than a hundred client AI SEO audits submitted to us over the past year or so, we've watched this exact scenario play out in enough variations that it needed to be documented. If you're an agency owner, an in-house marketer, or a founder trying to figure out what's real — this one's for you.

To be clear: AI tools aren't useless. Some are genuinely helpful for content gap analysis, keyword clustering, and first-draft briefs. But an SEO audit is a verification task. It requires crawling real URLs, pulling live data, and checking actual server responses. A large language model that can't access your site doesn't audit it. It generates a plausible-sounding document that borrows its structure from audits it has seen before. The difference matters enormously when you're deciding where to spend time and budget.

Pattern 1: Ranking Claims Backed by Nothing

This one shows up in nearly every AI-generated audit we've reviewed. A section titled something like "Current Keyword Performance" lists a set of target terms, assigns a ranking position to each, and flags the ones that need improvement. Looks authoritative. The problem? There's no data source.

No Google Search Console export. No Semrush or Ahrefs pull. No date range. No location context. When you ask clients where those numbers came from, they don't know. Sometimes the tool invents positions that are close enough to reality to feel credible — which is almost worse, because it erodes your ability to correct the record. Other times the numbers are completely fabricated for keywords the site barely touches.

I've seen ranking "data" in these documents that contradicts the client's own GSC dashboard by twenty or thirty positions on their most important keywords. When you're making decisions about where to invest in content or link building, a fabricated baseline is worse than no baseline at all. Real rank tracking requires a live connection to a data source — something LLMs simply don't have.

Pattern 2: Technical Errors That Don't Exist

The second pattern is the one that wastes the most developer time. A client's AI audit flags a page as returning a 404. The dev team investigates. The page returns a 200. Nobody touched it. The error was hallucinated.

This happens because an LLM has no ability to crawl a live URL and check its HTTP response code. It's pattern-matching against what SEO audits typically say. And since "broken internal links" and "crawl errors" are extremely common in real audits, the model includes them. For real audits, tools like Screaming Frog, Sitebulb, or a programmatic crawl via Ahrefs Site Audit or Google Search Console's Coverage report are non-negotiable. They hit actual server endpoints and return actual responses.

Other ghost errors we've seen in the wild: duplicate title tags reported for pages that have fully unique, well-crafted titles; missing canonical tags on pages that have them properly implemented; "thin content" flags on pages with three thousand words of substantive editorial. The audit reads as professional. The findings are fiction.

Pattern 3: Schema Recommendations That Break the Spec

This one is subtle enough that it fools experienced marketers — and it's the most technically dangerous of the three patterns. AI audits frequently recommend schema markup that either contradicts the current Schema.org specification or conflicts with Google's structured data guidelines. Sometimes both.

Common versions of this: recommending `Product` schema on a service page (Google's guidelines are explicit that Product is for physical or digital products, not services); adding `FAQPage` markup to pages where the questions aren't actually answered on that page; nesting `LocalBusiness` inside `Organization` in a way the validator flags as an error. Implementing any of these can result in a manual action for spammy structured data or, more quietly, just getting your rich results eligibility pulled.

I've audited sites that implemented schema from an AI recommendation and then went months without a rich result they'd previously earned. The markup looked right to a non-specialist. It failed Google's Rich Results Test. Always validate against the actual spec and test with real tools before you touch production markup.

Pattern 4: Page Speed "Fixes" With No Real Measurement

Almost every AI audit includes a Core Web Vitals section. Almost none of them pull real field data. There's a meaningful difference between lab data from a synthetic test and real-user data from the Chrome User Experience Report, which is what Google's PageSpeed Insights actually prioritizes for ranking purposes.

An AI audit might recommend "optimize your LCP" without checking whether your LCP is actually failing. It's a reasonable guess — a lot of sites do have LCP issues — but it's a guess. The actual issue might be a CLS problem caused by a single third-party script. Or the site might already be passing all three CWV signals. You don't know unless you measure. And that requires running the actual test on actual URLs.

Pattern 5: Content Advice Detached From Search Intent

AI audits love to recommend "adding more content" to underperforming pages. Which, fair enough — thin pages are a real problem. But here's where intent analysis matters and most LLM audits miss it entirely.

A page targeting a transactional keyword should be optimized for conversion, not padded with informational paragraphs. Adding a thousand-word explainer section to a product page can actually hurt its performance if it shifts the page's intent signal in Google's eyes. I've watched this play out on sites where well-meaning content additions moved pages from the top of page one to the bottom. More words is not the same as better words. Search intent analysis requires pulling the actual SERP, looking at what's ranking, and understanding what Google is already rewarding for that query.

Why This Keeps Happening

None of this is a knock on the engineers building LLM-based tools. The problem is structural. A language model generates text that is statistically consistent with its training data. SEO audits in that training data have certain predictable shapes: technical issues, content gaps, backlink analysis, schema recommendations. The model produces text that fits that shape. It has no mechanism to distinguish between "this error exists on this specific site" and "this error commonly appears in SEO audits."

And yes, this happens more than most agencies want to admit: clients bring these documents in, confident in the findings, and the agency either has to push back diplomatically or, worse, just starts working through the fake issue list to keep the client happy. Neither outcome is good for anyone.

What a Real Audit Actually Requires

A credible SEO audit has a few non-negotiable properties. It pulls live data. It verifies server responses. It connects to real ranking data with a source and a date range. And it cross-references recommendations against current specs and guidelines rather than pattern-matching against historical training data.

  • Crawl-verified technical findings — every flagged issue should trace back to a specific URL and an actual server response or HTML condition, confirmed by a crawl tool.
  • Sourced rank data — positions should come from a connected data source (GSC, Semrush, Ahrefs, or equivalent) with a date range and location context attached.
  • Schema validation against live tools — every structured data recommendation should pass Google's Rich Results Test and be checked against the current Schema.org spec.
  • Intent-aligned content recommendations — content gaps should be identified by analyzing what's actually ranking for the target query, not just by comparing word counts.
  • Real CWV measurement — Core Web Vitals assessments should use field data, not just lab estimates.

How Aergos Approaches This Differently

We built Aergos because we were tired of the same problem from the agency side. The platform actually crawls your site, pulls live rank data, and flags issues against verified criteria — not against a pattern-matched template. When you run a site audit through Aergos, the findings are tied to real crawl results: specific URLs, specific HTTP responses, specific on-page conditions. The schema checks cross-reference against current guidelines. The rank data comes from a live, sourced connection with date ranges and geographic context you can actually stand behind in a client meeting.

For agencies doing white-label SEO reports, that credibility matters every time you share a deliverable. A client who has already seen a hallucinated AI audit is going to ask hard questions. Your answers need to be backed by real data.

Where to Start If You've Inherited One of These Audits

Don't throw the whole document out. Some AI audits do surface real issues by coincidence, or they correctly identify broad strategic gaps even when the specific findings are wrong. The move is to verify, not dismiss wholesale.

  1. Take every technical claim to a crawl tool first. Run Screaming Frog or Sitebulb over the flagged URLs before you touch a single line of code. If the issue shows up in the crawl, it's real. If it doesn't, it's not.
  2. Pull your own rank data before you trust any position listed in the document. Connect GSC to your analytics setup or run a fresh Semrush/Ahrefs pull. Compare it to what the audit claims.
  3. Validate every schema recommendation in Google's Rich Results Test before implementation. If it fails the test or contradicts the current spec, it doesn't go into production.
  4. Run PageSpeed Insights on your actual URLs and look at the field data section, not just the lab score. Let that drive your CWV priority list.
  5. Check search intent before adding content. Pull the current top ten results for any keyword you're targeting and see what format and depth Google is already rewarding.

The bottom line is this: an AI audit is a document. A real audit is a process. If the tool generating it can't visit your URLs, it cannot audit your site. Full stop. Use LLM tools where they're genuinely strong — ideation, drafting, clustering — and use verification-based tools where accuracy is non-negotiable. Your clients are counting on you to know the difference.

Frequently Asked Questions

Matt Weitzman

About

Senior SEO Strategist & Co-Founder

Matt has over 15 years of experience in technical SEO and digital marketing. He specializes in algorithmic recovery, enterprise architecture, and leveraging AI for content scaling. He is a frequent speaker at search marketing conferences.

More articles by Matt Weitzman