A publisher's own site is the only thing a stranger is allowed to audit. This agent reads it the way a buyer would — where the email capture sits, what the ad load costs, whether the affiliate links still pay, what the schema qualifies for, what monetizes the betting traffic, whether the events take money — and returns a score with the evidence attached to every claim.
Homepage plus a handful of article pages, found through the sitemap rather than guessed. Robots respected, one request per second, every fetch cached so re-scoring never re-crawls. These are people worth emailing later — showing up as a bad crawler first is a cost, not a detail.
Each leak returns a severity, a confidence, and a revenue weight. The leak score is weighted severity across the checks that actually ran — and coverage reports how much of the model that was. A score at 71% coverage is a different claim from one at 100%, and saying so is the difference between an honest number and a flattering one.
Every finding carries the URL and the observation behind it. Not "newsletter capture could be stronger" but "no email input, ESP embed, or subscribe form on any of four content pages, here they are." A claim a stranger can verify in ten seconds is the only kind worth putting in front of one.
Weights are revenue impact for a regional or independent sports publisher. They are a defensible prior, not a measured finding — drawn from how these sites actually earn, and due to be re-derived from the benchmark once enough engagements show which fixes moved money. The page says so because the report says so.
The output goes to a stranger who did not ask for it. That makes a confident wrong answer more expensive than no answer, and most of the design work went into the places where the agent has to say it doesn't know.
Unscorable checks drop out of the denominator instead of quietly counting as clean. Scoring a check that failed to run would reward broken detection with a better number — exactly backwards.
Blocked egress, a refused user agent, and a genuine 404 are indistinguishable from outside. Only real HTTP failures are reported as dead; the rest are listed as unverified. Telling a publisher their links are broken when the network was the problem is the single most credibility-expensive mistake available.
Substack and its equivalents supply capture, schema, and a clean ad stack for free. Three checks then pass without the publisher choosing anything. Those runs are flagged, excluded from the benchmark, and reported as not comparable — otherwise they poison every percentile drawn from it afterward.
The leak weights are reasoned, not measured. The report says so in its own footer, and will keep saying so until the benchmark is large enough to replace them with evidence.
Twenty-seven fixture tests run without a network, each with a known-correct answer. They exist for one reason: to catch the false positive before it reaches somebody's inbox. Three found so far, all of which would have shipped.
One physical ad slot appears in the markup three times — as a defineSlot call, as a div with a matching id, and as an ad-ish class. Summing the matches tripled the count and flagged healthy sites as heavy. Fixed by counting two independent views of the same thing and taking the larger.
A raw operator link was being read as a betting unit, which closed the check and hid the worst version of the problem: betting traffic sent to the book with no affiliate parameter on it. Operator links now only count as monetization when they are actually tagged.
Found in a live end-to-end run, not in a test. Every affiliate probe returned no response and the report confidently called them broken. Now only real HTTP failures count as dead, and a site where nothing could be probed returns unknown rather than a flattering clean.
Every run appends one row — one per publisher, so re-scoring a site while tuning detectors never lets it vote twice. Percentiles stay suppressed until the sample is large enough to mean something, because a percentile drawn from four sites is a lie with a decimal point on it.
What accumulates is the part that cannot be copied by pointing a general-purpose model at the same page: a record of how mid-market sports publishers actually monetize, built one verified scan at a time. It is what turns "this could be better" into "this sits in the bottom third for sites your size," and it costs nothing extra per run.
The detectors are replaceable. Vendor fingerprints rot and get rewritten. The record does not.
Eight checks defined, six automated, twenty-seven fixture tests passing. Hosted-platform detection, evidence capture, and the benchmark are in. The first live runs against real publishers are pending — which is stated here rather than implied away, because fixture tests prove the logic works and prove nothing about the open web.
This agent measures one region of a larger map. The terrain it sits inside — every place a sports publisher earns, and what has to be true underneath each line — is documented separately.
The sports publisher monetization stackName the site and the report comes back: all eight leaks scored, the evidence behind each one, and the single finding worth acting on first. No charge, and nothing follows it unless you ask.
Every scan is run and read by hand before it goes out — an automated report that is confidently wrong about someone's own site is worse than no report. Expect a few days, not a few seconds. One email, no sequence, no list.