IndieDex Log 1: What am I doing with 960K pages

IndieDex has 960K programmatic pages, ~1K organic clicks a month, AI crawler storms that cost $15 a day, and no ads yet. A log of what I'm doing with it.

I want to log my thoughts on my project IndieDex. IndieDex started as a data toy: I just wanted to play around with a dataset, embedding models and AI. Then it became this huge thing that I wanted to make succeed. There is so much about it I wish I had done differently, but at this point I already invested a ton of money and time into IndieDex, and I want to let it be as it is and see what happens.

How I ended up with 960K pages

One of my fears, which I shared on my first post, is the fact that IndieDex has around 960K pages, which is way too many pages. I didn’t do it intentionally. I just knew I wanted a game details page and a games like page, and I wanted to offer the site translated so it could be accessible by as many people as I could. So I aimed to get translations working for all pages, and even for the generated content in the data, in French, Spanish and Portuguese together with English.

This means every 1 game gets 2 pages and every page gets 4 languages, so 1 game = 8 pages. Steam has around 120K games that are released (Valve’s own store search lists 181,539 games as of 2026-10-08, but many are unreleased), so 120K × 8 = 960K pages.

For scale: Google’s own crawl budget guide says it is written for “large sites (1 million+ unique pages)” and “medium or larger sites (10,000+ unique pages)”. IndieDex sits right under the first line.

I did all that before thinking about SEO. Once I launched and added it to Google Search Console, I was surprised at how fast it was getting clicks. It was like nothing I had ever seen. In the first month of launching it was already at nearly 1K clicks, with most of my pages being beyond the second page:

Google Search Console for indiedex.gg from Aug 9 to Aug 31, 2026: 914 clicks, 50.7K impressions, 1.8% CTR and an average position of 13.4, with clicks climbing through the month

The first month in Search Console. Click the image for full size.

Social flopped, SEO didn’t

Initially I was thinking IndieDex would spread socially, with people wanting to make lists and share their profiles. I was very wrong. There have been people that created profiles and some created lists, but nobody really seemed to care about those features. I even had a feature to match people together based on their tastes in games. I ended up turning that off later, since nobody was using it and it is complex to maintain.

So after a few months of seeing the traffic grow, that piqued my interest in SEO again. Then September was flat, and I thought “that is it, it’s going to die”. So I made a few changes, like making each page more valuable: the game details page now shows users what they would like about the game and what they may not like about the game, based on my own unique data tagging and analysis:

IndieDex game details page: an "Is it for you?" section with "You'll like it if" and "Skip it if" lists, and "What players are saying" with pros and cons based on Steam reviews

That went live in late September, maybe the 29th. The GSC bar moved right after, on Oct 2nd, but I don’t think it was related to my changes. I believe it was related to Google’s September update: the August update had caused my site to go flat, then this September one hit and it spiked. I shared that on my first post.

Per Search Engine Journal, Google’s September 2026 spam update launched on September 24 and the Search Status Dashboard marked it complete on October 8, 2026. Google hasn’t said what this one targeted.

I have received great comments and feedback about how much people like the search engine and the content added to the pages. You can also like/dislike and leave reviews, although very few people are engaging with those features. I am leaving it there.

They also enjoy the games like page, which are the most popular SEO pages.

IndieDex "Games like Dark Souls: Remastered" page: what the games share, why these picks, a "start with" list and three recommended games with tags, review scores and play time

The search engine is my masterpiece, and improving it costs too much

It took me a lot of effort to build IndieDex. I remember going through at least 4 iterations of the search engine design, and countless iterations testing it until I was happy with it. That search feature is my masterpiece, and I know it can get a lot better. But working with such a large dataset and 4 languages makes things a lot more expensive.

One thing I want to add is processing the game images, to be able to add the style of the game as a tag from the image, instead of relying on it being described by one of the data sources. The plan was to use EmbeddingGemma 2 and send about 3 images per game for 120K games, to make it more accurate. Unfortunately, I ran the cost calculations for that and it will cost me somewhere around $15,000 to get it done properly. For a project that has no ads or subscriptions, that is too much for me.

That’s EmbeddingGemma 2 on the model card: a 740M-parameter open embedding model (Apache 2.0) that “unifies 4 modalities (text, images, video, and audio)”, so one model could embed the screenshots next to the text fingerprints. Three images per game across 120K games is about 360K image embeddings.

So I’m going to stop touching it

So I decided: I’m going to stop touching it for a while, and I’m going to see where it goes from here on its own.

With my concerns of Google cracking down on it due to the large number of pages, I set a programmatic noindex on some pages that don’t have as much data as others. If a game is missing social signals and doesn’t have enough system tags on the data, I ignore it, as it would be a lower accuracy judgement over what the game is and what people experience with the game. Only about 40% of the pages will get indexed currently, and even that is too much.

Also, choosing to not completely hide the noindex pages means Google will still spend crawl budget on them. Just something I gotta live with. Google’s crawl budget guide confirms the trade-off: for pages you don’t want crawled at all it says “Don’t use noindex, as Google will still request, but then drop the page”, and points to robots.txt instead.

I do worry about newer games too. Those wouldn’t have as much data from the release, so every so often I might run a backfill job to ensure everything gets a fair shot at being indexed, but also gets as much data as it can have, to provide as much value as possible to the end user.

The closest public data point on which programmatic pages Google keeps comes from a much smaller experiment, 225 pages with almost no backlinks:

“Hub/category pages indexed at 87%, leaf pages at 18% … The biggest lesson: Google doesn’t mind programmatic content if it doesn’t look programmatic. The pages that got indexed were the ones where the content happened to diverge most from the template pattern.”

— Arnjen on Hacker News

What it costs to keep alive

So far, I have spent over $1000 processing the data for the IndieDex catalog. And there have also been a lot of challenges from the infrastructure side: such a large footprint has caused AI crawlers to go crazy on my site and make a ridiculous amount of requests, all of which cost me money in Neon DB staying up and Vercel compute being used. I had a lot of crawling from ClaudeBot and AWS.

Vercel Pro plan daily usage for IndieDex from Aug 3 to Oct 8, 2026: most days under $2, with spikes to around $15 a day in early September and another around September 25

Vercel daily consumption. The tall bars are crawl storms.

ClaudeBot showing up first is not a coincidence. In Cloudflare’s Radar 2025 Year in Review, Anthropic had the highest crawl-to-refer ratio of any AI company: as much as 500,000 pages crawled per visitor referred at the peak, and between about 25,000:1 and 100,000:1 after May 2025. Google’s ratio was 3:1 by mid-July. And in June 2026, Cloudflare’s CEO reported that bots had passed humans at 57.5% of HTML requests on their network, “faster than I predicted”.

And so many bots and scrapers keep hitting it. I set up the firewall with rate limiting, but I always feel like I’m playing whack-a-mole with these, and sometimes I will overdo it and users will complain about 403s and 429s. Normally I see the usage cost me anywhere from $0.50 to $2.00 per day, but when crawl storms start I can see it cost me as much as $15.00 per day, and if I didn’t stop it, it would be even more expensive.

Same stack, same problem. An engineer at Metacast, after a crawler-driven Vercel bill:

“Our biggest issue right now is unidentified crawlers with user agents resembling regular users. We get hundreds of thousands of requests from those daily and I’m not sure how to block them on Vercel.”

— ilyabez on Hacker News

Why rate limiting feels like whack-a-mole, from a thread on AI crawler defenses:

“AI crawlers do very little requests from an ip address to bypass rate-limiting. Last year they could still be blocked by ip range, but now the requests are from so many different networks that doesn’t work anymore.”

— tortillasauce on Hacker News

ISR is the bet

Now I have a long term cache, so pages get cached for months. I can’t generate 960K pages statically at build time, so I keep it ISR.

ISR (Incremental Static Regeneration) means a page is rendered the first time someone asks for it, then served from Vercel’s cache until it expires or is revalidated; the docs describe it as “a caching strategy that combines the speed of static content with the flexibility of server-side rendering”. Cached pages stay until they go “unaccessed for 31 days”, and a revalidation that produces identical content costs no write units.

I still get surprised though. Like today, the ISR writes went up a lot, and I thought it was another crawl storm or something. But I had also posted on r/InternetIsBeautiful, and it did get kind of popular, so what I thought were bots might actually have been a ton of users. So far the post is at 63 likes, 51K views and 8 comments, with 92 shares. It has provided a nice traffic surge that helped me identify some issues with my pre-fetching strategy. I am now making some improvements to that, but I hope this will be my last deploy for a few months.

The bankruptcy-by-virality fear has a counterexample on the same stack. After a site went viral on X in December 2025:

“After millions of unique visitors we racked up about $10 in costs. … Vercel + Next is cheap if cached correctly.”

— lukeigel on Hacker News

No monetization yet

And that really worries me, because I don’t have monetization set up at all. I know I’ll be able to use ads on it, but for Mediavine Journey you need a site that does over 1K in traffic per month. I tried applying to it last month when my traffic hit 1K, but they denied me, so I gotta wait. I just hope I don’t get too much traffic by then, it would cost me a lot of money…

The denial just said I didn’t qualify, no specifics. Reading Mediavine’s application requirements now, it says 1K traffic from Tier 1 countries, so that was likely the problem: most of my traffic is US, but a lot of it is South America too. So I probably didn’t have 1K Tier 1, and that is what blocked me. I’m pretty far from 1K Tier 1:

Search Console countries report for indiedex.gg: Brazil leads with 463 clicks, then United States 326, France 213, India 168, Spain 136, Indonesia 68, Philippines 67, Mexico 66, Iran 59 and United Kingdom 58

Clicks by country in Search Console. Brazil is ahead of the US.

Mediavine’s requirements for Journey: “Receive a minimum of 1,000 sessions from Tier 1 countries (including the U.S., Canada, the U.K., and Australia) within a 30-day period”, original audience-first content, and their content quality and traffic standards. No site age requirement is listed. The full Mediavine tier needs “a minimum of $5,000 in annual ad revenue”.

So what is the plan now?

Wait a few more months, keep monitoring Vercel, PostHog and GSC, and in early December I will be applying to Mediavine again. Hopefully my traffic won’t surprise me, or at least, I hope the ISR protects me from high costs. I built this to be as affordable to run on the long-term as possible, that way it has the best chance to reach a wide audience.

I really hope Google doesn’t kill it though. The SEO is a lot more effective than the social campaign, and it is a lot less effort for me with a full time job and tons of other sites I need to maintain.

I’ll do a log 2 post when December comes and I apply again, or if I get a ridiculous spike in traffic and go bankrupt.

§ 1/8 · How I ended up with 960K pages0%