Back to all posts

TIME’s Secret Markdown Layer for AI Crawlers

These elements are invisible to human readers but are deliberately structured to be processed by large language models as ad impressions.

AI CrawlersMachine LearningWeb ScrapingLLM Data
main thumbnail for TIME’s Secret Markdown Layer for AI Crawlers
main thumbnail for TIME’s Secret Markdown Layer for AI Crawlers
Reader Lens

Automation needs a narrow first win

The best first AI workflow is usually a repeated task with a clear input, clear output, and a human approval step.

TIME is officially bifurcating the user experience into two distinct tracks: one for humans and a stripped-down markdown version designed specifically for AI crawlers. By identifying requesting bots via User-Agent headers, the publication serves different content formats to different audiences. While humans see a standard HTML layout, crawlers like ClaudeBot, PerplexityBot, and OAI-SearchBot are served markdown files that are significantly smaller—roughly 13,409 bytes compared to the 303,235 bytes of the original HTML.

The Architecture of Machine-Readable Ads

This isn't just a move toward better "readability" for models; it is a calculated infrastructure for machine-based advertising. The markdown versions include embedded sponsored content and ads from brands like Ally Bank and the Project Management Institute. These elements are invisible to human readers but are deliberately structured to be processed by large language models as ad impressions.

Crucially, the tracking mechanism for these ads has shifted. Instead of traditional pageviews, TIME is tracking impressions based on the number of tokens fed into the models. This represents a fundamental shift in how "reach" is measured in an era where the primary consumer of content is no longer a person clicking a link, but a model processing a data stream. Furthermore, the gatekeeping is surgical: while OAI-SearchBot is allowed in, GPTBot and ChatGPT-User are met with 406 errors. This suggests TIME is being highly specific about which models get to see this curated, ad-supported data stream.

The Rise of the Two-Tier Internet

The part worth being skeptical of here isn't the existence of the markdown—it's the implications of what this means for the integrity of the "open" web. When bot traffic already outnumbers human traffic on most days, the primary audience for a publication like TIME is no longer a person, but an algorithm. By creating a layer of the site written entirely for machines that humans have no idea exists, TIME is essentially creating a "shadow web.

In practice, we are moving toward a reality where the information provided to an AI assistant is a curated, monetized version of reality that is decoupled from the public-facing site. The real story here isn't just a new way to serve ads; it's the formalization of a two-tier internet. One tier is for human consumption, and the other is a data pipeline where "truth" is filtered through the requirements of token efficiency and advertiser placement. This creates a significant information asymmetry: what the AI tells you might not be what a human would see if they visited the same URL. If the goal of AI is to summarize the world, we need to be aware that the "summary" is increasingly being served from a source that prioritizes machine-readability and ad-token counts over a unified experience. This points to a future where the "source material" for AI is no longer a reflection of the public record, but a specifically engineered product designed for consumption by models.

inside paper visual for TIME’s Secret Markdown Layer for AI Crawlers
main thumbnail for TIME’s Secret Markdown Layer for AI Crawlers
closing highlight visual for TIME’s Secret Markdown Layer for AI Crawlers
main thumbnail for TIME’s Secret Markdown Layer for AI Crawlers
Source and trust note

Built from source research and filtered through practical implementation judgment.

Reference: www.vincentschmalbach.com

Got a question about how this applies to you? →

Keep reading

Follow the thread