TL;DR: Bing had discovered around 400 URLs from our sitemaps and crawled almost none of them. After ruling out verification, on-page, and discovery problems one by one, the real cause was crawl-budget starvation: a young domain with 3 backlinks does not earn crawl budget. So the work moved off-site to authority, while we shipped every fix that was genuinely worth shipping and declined two that were not. The machine plane also became a daily series this week, running 64 to 194 AI-engine requests a day from Sep 17 to 20. Search Console for the week: 8 clicks, 811 impressions, 172 pages.
This site is two surfaces. One is a marketing site for humans. The other is a machine-first registry of open-source GTM tools, agent-operated and human-gated, updated daily, published with structured data and plain-text twins so AI engines can read it as easily as people can. It went public this month with essentially zero search presence, so this is a from-zero visibility program reported in the open.
Why Bing is the story this week
ChatGPT and Copilot answer from Bing's index. That is public knowledge, and it changes the stakes. Zero presence in Bing does not just cost a slice of search traffic; it makes the site invisible to a large share of AI answers. So the week opened with a direct question: what is actually in Bing's index?
The numbers were stark. Bing had discovered roughly 400 URLs from the submitted sitemaps and crawled essentially none of them. Three URLs total sat in its index, and not one was a registry page. I ran URL inspection on 7 priority pages by hand. All 7 came back the same way: "Discovered but not crawled."
That phrase is the whole diagnosis in three words, if you are willing to read it honestly.
Ruling causes out, one at a time
The tempting move is to assume it is your fault in code and start rewriting. I did the opposite and tried to falsify each cheap explanation first.
- Verification: confirmed, the property is verified. Not the problem.
- On-page: the priority pages are clean, well-formed, and indexable. Not the problem.
- Discovery: Bing had already discovered the URLs from the sitemaps. Not the problem.
When every technical cause checks out and the pages still are not crawled, the remaining explanation is the unflattering one. This is crawl-budget starvation. A young domain with 3 backlinks does not earn enough crawl budget for a search engine to spend on 400 pages. The binding constraint is authority, not code. The domain is young even though the practice behind it is not; I have been delivering growth systems under contract since August 2024. Search engines do not read your track record, only your link graph, and this link graph is a few weeks old.
Naming that plainly matters more than any single fix, because it tells you where the work actually is.
What shipped anyway
Naming the constraint does not excuse skipping the real fixes. Several were worth doing on their own merits, so they shipped:
- A bulk IndexNow re-seed of 403 URLs, accepted. Bing recorded 520 URL submissions in 17 hours.
- The 7 priority pages queued manually for indexing.
- Social-preview tags added at the registry template level, so one change fanned out to around 400 pages at once.
- Meta descriptions trimmed to spec, including the homepage from 231 down to 157 characters.
- Internal links cleaned to their canonical form so crawlers stop paying for redirect hops.
Each of these is a genuine improvement regardless of the crawl-budget verdict. None of them, on their own, buys authority.
What I refused to build
Two proposals came up that sounded reasonable and were the wrong spend. Saying no is part of the report.
The first was a sitemap-index restructure to fix a supposed URL gap. The generator already handles indexing correctly, and the "gap" was a miscount. The live URL count is 403, verified. There was nothing to fix, so nothing was built.
The second was a redirect fix for a legacy page. It has zero backlinks to recover and lives in shared infrastructure where a change carries real blast radius. Cost and benefit pointed the same direction: leave it alone.
Both decisions cost nothing and protected the week from busywork dressed up as progress.
Moving the work off-site
If the constraint is authority, the work is not in the codebase. So the response was to raise the backlink ceiling: directory and community submissions, several live and several pending. I also wired the organization's social profiles (YouTube, Instagram, TikTok) into the site's structured data, so machines can corroborate that the entity is real.
Some channels were deliberately deferred rather than forced. Paid-only listings are parked. Community posts wait until the account can carry them without looking like drive-by link drops. When the constraint is authority, forcing distribution early burns the channel you will want later.
One measurement milestone landed alongside this. A server-side pipeline now feeds CDN logs into product analytics on a daily schedule, so agent traffic (AI crawlers and answer-engine fetchers) is a first-class dashboard rather than a monthly spelunking session. The machine plane finally gets watched like the human plane.
The machine plane, now day by day
Last week I could only report a single two-day snapshot, because logging had just started on Sep 11 and the daily pipeline did not exist yet. That partial baseline read 322 AI-engine requests across the first two days, from every major answer engine. This week the daily pipeline turned that one-off dig into a clean per-day series, so I can actually show the machine plane moving.
One honest gap first: the daily series only becomes complete from Sep 17 onward. Sep 14, 15 and 16 were the days the pipeline was still being built and backfilled, and I do not have a clean per-day split for them, so I am not going to invent one. What follows is what the daily series actually holds.
AI answer-engine requests (OpenAI plus Anthropic plus Perplexity), per UTC day:
- Sep 17: 84 AI-engine requests (746 human), 882 total.
- Sep 18: 64 AI-engine requests (583 human), 676 total.
- Sep 19: 194 AI-engine requests (516 human), 846 total.
- Sep 20: 118 AI-engine requests (1,558 human), 1,716 total.
The named engines behind those counts are the same ones from the first snapshot: OpenAI (GPTBot, SearchBot, ChatGPT-User), Anthropic (ClaudeBot, Claude-User) and Perplexity, alongside the traditional search crawlers Googlebot and bingbot tracked separately. Anthropic's crawlers were the steadiest presence, running 42 to 50 requests a day; OpenAI spiked to 141 on Sep 19 and settled again. Small numbers, but they are daily now, and daily is the whole point.
This is the same series that drives the traffic chart on the site's homepage, broken out by audience and refreshed every day, so anyone can watch the machine plane move in the open rather than take my word for it here.
The enrichment wave and the numbers
The second search-driven enrichment wave went out on Sep 15: 13 more registry pages, chosen because search or AI surfaces already show them (7 had Google generative-feature impressions). One more license was corrected on re-verification, Apache to AGPL. Accuracy is the product, so a correction is a feature, not an embarrassment.
The trajectory numbers, all small, all real, all moving:
- Search Console now reads 8 clicks, 811 impressions, 172 pages, against 6 / 389 / 127 the week before.
- Daily impressions ran 18, 174, 197, 222, 175 across Sep 8 to 12.
- Generative-AI-feature impressions total 15 across 11 pages, and the top query sits around position 5 with 25 impressions.
What the week bought
The constraint is named: authority. The machinery is built: an indexing push, a now-daily agent-plane measurement running 64 to 194 AI-engine requests a day on the homepage chart, and a repeatable enrichment loop. Next is compounding, not cleverness: more proof-driven enrichment waves, more legitimate backlinks, and a patient watch on whether Bing starts spending crawl budget once the domain earns it.