
Most agents pass their evals and fail in production. Prefactor is the evaluation layer that closes the gap. We score every agent run in real time, surface quality regressions and drift as they happen, and show engineering teams exactly how their agents are performing at scale. Built for the teams shipping agents to customers.
Prefactor is a SaaS tool that evaluates AI agents in real-time, identifying quality regressions and performance issues as they occur. It provides engineering teams with insights into agent performance at scale, ensuring reliability before deployment.
Overall, the comments reflect strong interest and excitement about Prefactor's capabilities, with some concerns about evaluation reliability.
<p>Hi Product Hunt - Matt here, co-founder of Prefactor with Simon. </p><p></p><p>We've been heads-down on this one for a while, so finally getting to show you is a real thrill.</p><p></p><p>Let me start with the question this whole thing is built around: <strong>do you actually know what your agents are doing in production right now?</strong></p><p></p><p>For most teams we talk to, the honest answer is "…not really." Your evals pass, everything's <em>green</em>, you ship - and then every real run vanishes into a <strong>black box</strong>. Quality quietly drifts. Risk creeps in. Costs climb. And eventually someone asks which of your agents are still doing their job, and the room goes quiet.</p><p></p><p>That silence is the entire reason Prefactor exists. Gartner reckons <strong><em>40%+ of agentic AI projects get scrapped by 2027</em></strong> - and honestly, this gap is a big part of why.</p><p></p><p>So we built the thing we kept wishing we had: Prefactor scores every run in production the moment it happens - quality, drift, risk - then wires those scores straight into action. A failing agent gets caught live, not charted three days later.</p><p></p><p><strong>How it actually works</strong></p><ol><li><p><strong><em>prefactor init </em></strong>- one command connects your workspace and discovers your agents across your runtimes. First traced run in under 5 minutes.</p></li><li><p><strong><em>Drop in the SDK </em></strong>(TypeScript or Python) - native for LangChain, Claude, Vercel AI, OpenClaw and LiveKit. Every call becomes a span, streaming in live with cost and data risk attached.</p></li><li><p><strong><em>Run the evals </em></strong>you define on every run - LLM-as-judge, technical checks, qualitative metrics. Custom spans pull context from GitHub, Linear, Jira or your database, so every eval is grounded in what actually happened, not a guess.</p></li><li><p><strong>Act.</strong> Hold, approve or block the second a run crosses a line - automatically at runtime, or routed to a human. Every decision logged and enforced through the SDK or API.</p></li></ol><p>The payoff: you get to ship agents like real software — versioned, staged, promoted through dev → staging → prod only when evals pass, with instant rollback when they don't.</p><p></p><p><strong>Why it's different</strong></p><p>Most tools observe and score, then hand you the problem. Prefactor closes the loop - observe, evaluate, act, all inside the same run. A risky agent gets caught, not just charted.</p><p>My favourite bit of feedback so far: a customer with 40 agents in production and, in their words, "no honest way to say which ones were still doing their job." We gave them that answer - and the brake pedal for when one wasn't.</p><p></p><p><strong>Who it's for</strong></p><p>Engineering teams shipping agents to real customers, on any stack. Agent frameworks work natively; everything else plugs in through OpenTelemetry or the core SDK.</p><p></p><p><strong>A little something for the PH community</strong></p><p>Sign up today and you get <strong>1,000,000 free agent steps</strong>. Every span counts, so that's a serious amount of live production evaluation on us. Valid until Friday 11:59pm PT, once you've set up your first agent.</p><p></p><p><strong>Our ask</strong></p><p>If you're running agents in production, tell us how you keep tabs on them today - even if the honest answer is "we're mostly hoping." We'll be in the comments all day and we'd genuinely love to hear what's working, what's breaking, and what you'd want a tool like this to do next.</p><p></p><p><strong><em>Get started free at </em></strong><a href="http://prefactor.tech" target="_blank" rel="nofollow noopener noreferrer"><strong><em>prefactor.tech</em></strong></a><strong><em>.</em></strong> First 25,000 spans a month free, no card needed. </p><p></p><p><a href="https://www.producthunt.com/@simon_russell1" target="_blank" rel="nofollow noopener noreferrer">@simon_russell1</a> <a href="https://ww
<p>Congrats Ethan and team!</p>
<p>Genuinely curious what "scoring every run" looks like at scale, is that sampled or are you actually evaluating 100% of production traffic? That distinction matters a lot for cost and for trust in the numbers.</p>
<p>Drift is the one nobody talks about until it's already cost them a bad week. Real detection here feels like the right instinct.</p>