
One API to scrape, enrich, and extract the internet
Context.dev is the web context API for AI products and agents. Scrape any URL, crawl sites, turn pages into LLM-ready Markdown, extract structured data into your own schema, capture screenshots, and retrieve logos, colors, fonts, styleguides, company data, and transaction enrichment through one API. YC-backed, no card required, and built so developers or coding agents can integrate in minutes.
Context.dev is an API that enables users to scrape, enrich, and extract data from the internet, converting web pages into LLM-ready Markdown and structured data. It offers features such as screenshot capture and retrieval of company assets, designed for easy integration by developers.
Overall, commenters express strong satisfaction with Context.dev's capabilities and support, while raising some concerns about pricing and specific technical challenges.
Hey Product Hunt 👋 I’m Yahia, founder of Context.dev. I built Context.dev because every AI product eventually runs into the same problem: models are powerful, but they don’t know what’s happening on the live web. So teams end up building the same annoying infrastructure over and over again: scrapers, crawlers, browser rendering, proxy handling, sitemap parsing, Markdown cleanup, screenshots, logo extraction, brand enrichment, company data pipelines, and more. Context.dev turns all of that into one API. You can scrape any URL, crawl a site, extract clean LLM-ready Markdown, pull structured data into your own schema, capture screenshots, retrieve logos/colors/fonts/styleguides, enrich companies, and give your agents fresh web context in seconds. The part I’m most excited about: Context.dev is agent-native. You can integrate it yourself, or paste one line into your coding agent and let it sign up, grab an API key, and wire the API into your codebase. We’re YC-backed, have a free tier with no card required, and are already powering products at teams like Mintlify, daily.dev, DocsBot, Chatwoot, and more. Would genuinely love feedback from the PH community, especially from anyone building AI agents, RAG pipelines, onboarding flows, enrichment workflows, or anything that needs live web data. Happy to answer questions all day!
<p>Right, I wasn't doubting the render fidelity, I meant determinism across fetches. Same URL scraped today vs next week: if the live DOM reorders a section, the markdown shape moves with it and an agent that indexed against the first shape drifts. Do you expose a content hash or a diff between fetches, so a pipeline can tell 'page actually changed' from 'page just reordered'? That's the bit that decides whether I wire it into an agent loop or keep it a one-off pull.</p>
<p>Pricing tied to successful scrapes only looks useful! Paying for failed fetches could be painful, I know this from own experience. Curious on the stealth layer as plenty of sites serve completely different content by visitor country (price, availability, language)... Can I pin the exit region per request or does the geo just fall out of whatever proxy the pool grabs that day?</p>
<p>We are using <a href="https://Context.dev" target="_blank" rel="nofollow noopener noreferrer">Context.dev</a> and we love it! Recommended</p>