On April 16, 2026, Factory AI announced a $150M Series C at a $1.5B post-money valuation. Khosla Ventures led the round, Sequoia Capital co-invested, and Keith Rabois joined the board. Three years ago, Matan Grinberg was a UC Berkeley PhD student cold-emailing investors. Today, Morgan Stanley and Ernst & Young are paying for his product. That’s a story worth understanding in detail.
We’ve been watching Factory since their Series A. What we find most interesting isn’t the valuation — it’s what the product actually does inside enterprise engineering teams, and what the bet on dedicated AI coding agents says about where software development is heading. You should care about this whether you’re a CTO, an engineer, or someone who manages people who are.
What Factory Actually Does — and Why Enterprises Are Paying
Factory builds AI agents it calls Droids. These aren’t copilots sitting beside a developer in an IDE — they’re autonomous agents designed to handle multi-step engineering tasks end to end. Pull request reviews, test generation, codebase migrations, bug triage, documentation. Work that developers do, not suggestions developers act on.
The model-agnostic architecture is the technical differentiator Factory keeps returning to. Droid doesn’t depend on a single foundation model. It can route tasks through Anthropic’s Claude, OpenAI’s models, or DeepSeek depending on what the task demands. That flexibility matters more than it might seem: no single model is best at everything, and enterprise customers don’t want to be locked into one AI vendor’s capabilities or pricing.
Morgan Stanley is using Factory Droids for engineering automation. EY has integrated them into enterprise development workflows. Palo Alto Networks — a cybersecurity company with some of the most stringent code quality requirements in the industry — is running them too. These aren’t pilot programs at experimental startups. These are serious, risk-conscious organizations that have moved past evaluation into actual deployment.
The Numbers Behind the $1.5B Valuation
Let’s be precise about the funding trajectory. Factory raised a $15M Series A from Sequoia. Then a $50M Series B at a $300M valuation. The Series C at $1.5B represents a 5x valuation jump from Series B — in what we estimate is under 18 months between rounds.
The benchmark numbers are worth examining carefully. On Terminal-Bench v0.1.1 — 80 human-verified, Dockerized engineering tasks — Factory’s Droid scored 58.75% using Claude Opus 4.1. On Terminal-Bench v2.0 (89 tasks), Droid with GPT-5.3-Codex scored 77.3%, ranking second overall behind Forge at 78.4%. Factory’s agents occupy three of the top five slots on that leaderboard.
For comparison: Droid outperforms Codex CLI (42.8%) in head-to-head testing and beats earlier Claude-based competitor configurations at 43.2%. These aren’t marketing numbers — Terminal-Bench is an independent benchmark with Dockerized task verification. You can review the methodology yourself at tbench.ai.
Blackstone and Insight Partners also participated in the Series C, alongside earlier backers including NEA, 20VC, and Mantis VC. When growth equity firms like Blackstone enter an AI developer tools deal, it signals revenue — not just product promise.
The Risk Nobody’s Talking About Loudly
We want to be direct here: the biggest risk in autonomous AI coding agents isn’t the technology failing. It’s the technology succeeding in ways engineering teams haven’t prepared for.
When an agent writes and merges code autonomously, accountability becomes complicated. If Droid generates a security vulnerability in a Palo Alto Networks codebase — or a compliance issue in a Morgan Stanley trading system — who owns that? The vendor? The team that deployed it? The engineer who approved the PR? These questions don’t have clean answers yet, and the enterprise contracts being signed today are going to be tested by the first serious incident.
There’s also a skills question we find ourselves thinking about more than most coverage addresses. If your junior engineers stop doing the work that makes them senior engineers — debugging, reading unfamiliar codebases, writing tests from scratch — what happens to your team’s capability over a three-year horizon? We’re not anti-automation. We use AI tools aggressively in our own engineering work. But we think about this deliberately, and you should too.
Factory’s model-agnostic architecture also cuts both ways as a risk. Flexibility is a strength, but it also means Factory’s competitive moat isn’t a proprietary model — it’s orchestration, workflow design, and enterprise integration. Those are defensible, but they’re also more imitable than a trained foundation model. As Anthropic, Cursor, and Cognition all push deeper into agentic workflows, Factory will need to keep moving.
3 Questions Every CTO Should Ask Before Buying AI Coding Agents
We’ve worked with enough engineering organizations to know that the evaluation questions most teams are asking about AI coding tools are too shallow. Here’s what we’d push you to dig into.
1. Where does human review actually happen? Not in the sales deck — in the actual workflow. Which tasks does the agent complete autonomously, which require approval, and who sets those thresholds? If the answer is vague, that’s a red flag. You need a clear human-in-the-loop map before anything touches production.
2. How does the vendor handle agent errors in regulated environments? If you’re in financial services, healthcare, or cybersecurity, an AI-generated compliance failure isn’t just a bug — it’s a regulatory event. Ask Factory (or any vendor) for documented incident response procedures specific to autonomous code generation. If they don’t have one, they’re not ready for your environment.
3. What does success look like at 12 months, not 90 days? Pilot programs for AI coding tools almost always show productivity gains in the first quarter. The harder question is whether your team’s overall engineering capability — including the humans — is stronger or weaker a year in. Build that measurement into your evaluation criteria from day one.
What This Means for Software Teams Right Now
The Factory Series C is a signal, not an anomaly. AI coding is, as the market keeps confirming, the most commercially validated application of generative AI in enterprise software. The question for your team isn’t whether to engage with these tools — that decision is already being made for you by your competitors. The question is how deliberately you engage.
We think the teams that win with AI coding agents in the next two years will be the ones who treat deployment as an organizational design problem, not a procurement problem. You’re not buying software — you’re restructuring how engineering work gets allocated between humans and agents. That requires intentional decisions about where agents operate freely, where humans stay in the loop, and how you measure outcomes that matter beyond lines of code shipped.
Factory’s $1.5B valuation tells you investors believe the autonomous agent approach beats the copilot approach at scale. The enterprise customer list tells you at least some of the most sophisticated engineering organizations in the world agree. Whether that’s right for your team depends on factors no benchmark can answer for you.
At ExaEdge, we build with AI tools and we think carefully about how they change team dynamics — not just productivity metrics. If you’re evaluating AI coding agents for your engineering organization and want a direct conversation about what we’ve seen work and what hasn’t, get in touch with us. You can also read our broader take on AI infrastructure investment trends in the ExaEdge blog, and see the kinds of engineering problems we tackle in our own projects.



