AI Can Write Code. Why Isn’t Software Better?
AI가 코드를 작성할 수 있게 된 지금, 왜 소프트웨어는 실질적으로 나아지지 않았는가? 이 질문이 TypeSafe AI의 Diogo Almeida가 Jev를 만든 근본 동기이며, a16z의 Ben Horowitz, Martin Casado와의 이번 대화는 현재 AI 투자 내러티브의 가장 큰 맹점을 정면으로 겨냥합니다.
Almeida의 핵심 주장은 간명합니다. Claude Code, Codex, Cursor 같은 코딩 에이전트는 소프트웨어를 더 빠르게 생산하지만, 10년 전과 동일한 종류의 소프트웨어를 만들 뿐입니다. 진짜 문제는 속도가 아니라 소프트웨어 자체의 표현력(expressive power)입니다. Jev는 이를 해결하기 위한 새로운 프로그래밍 원시 타입(primitive)으로, 개발자가 자연어 의도를 상태 머신의 결정으로 변환할 수 있게 합니다. 즉, 소프트웨어 엔지니어링을 자동화하는 것이 아니라, 소프트웨어 자체가 할 수 있는 일을 확대하는 것입니다. Almeida는 이를 1970년대에 사실상 소멸된 확률적 프로그래밍(probabilistic programming)의 부활이자, 기존 로직 게이트에 '작은 두뇌'를 추가하는 것으로 비유합니다.
근거는 설득력 있습니다. OpenAI는 2020년부터 고객 서비스 자동화를 시도했지만 여전히 실현되지 않았습니다. RLHF의 일반화 능력이 입증된 이후에도, AI 산업은 인간 평가자를 최적화하는 방향으로 흘러가며 실질적 작업 자동화와는 멀어졌습니다. Almeida가 말하는 '신뢰성'은 단순한 업타임 SLA가 아니라, 개발자가 예시 쿼리 없이도 Jev를 신뢰하고 프로그래밍할 수 있는 수준의 지능적 일관성을 의미합니다.
함의는 SaaS 시장을 포함해 광범위합니다. 코딩 에이전트 등장 시 폭락했던 SaaS 기업 가치가 Jev 출시 이후 오히려 반등한 것은, 이 기술이 기존 소프트웨어를 대체하는 것이 아니라 극적으로 강화할 수 있음을 시사합니다. 수백만 명의 기존 사용자에게 이미 도달한 SaaS 기업들이, Jev 같은 지능형 레이어를 통해 소프트웨어 경험 자체를 재정의할 수 있다면, 이는 단순한 기능 추가가 아닌 플랫폼 재편에 가깝습니다.
Coding agents like Cursor and Codex have generated enormous enthusiasm, but Diogo Almeida — founder of TypeSafe AI and former OpenAI researcher — argues they are solving the wrong problem. Producing code faster is not the same as producing better software, and the distinction matters enormously for where AI's real economic value will be captured.
Almeida's central thesis is that software itself needs to become smarter, not merely faster to write. Jev, TypeSafe's core product, is a new programming primitive that maps natural language intent to state-machine decisions with confidence levels — something closer to a classifier embedded in code than a chatbot bolted onto an app. The conceptual shift is significant: instead of automating the software engineer, Jev expands what software can express and do autonomously. Almeida frames this as a revival of probabilistic programming, a field that largely died in the 1970s, now made practical by modern language models.
The supporting evidence is pointed. OpenAI has been attempting to automate customer service since 2020 without success. GPQA benchmark scores suggest models can answer doctoral-level questions, yet a fast-food drive-through remains unautomated. Almeida's diagnosis: after RLHF demonstrated remarkable generalization in late 2021, the industry optimized for human evaluators rather than genuine task automation — overpromising on capability while underdelivering on reliability. His North Star metric is 'intelligence per dollar,' not benchmark performance.
The implications for SaaS companies are counterintuitive. When coding agents emerged, markets priced SaaS for disruption; when Jev launched, the same companies celebrated. If intelligence becomes an embeddable primitive rather than a replacement product, incumbents with existing user distribution gain a compounding advantage — they need only make their software dramatically more capable, not rebuild from scratch. The deeper structural bet is that AI's most durable economic contribution will flow through existing software, not around it.
코딩 에이전트는 더 빠른 삽질일 뿐 — 소프트웨어의 표현력은 그대로입니다
Claude Code, Codex, Cursor는 10년 전 인간이 작성하던 것과 동일한 종류의 코드를 더 빠르게 생산합니다. Almeida는 이를 '속도 향상'이 아닌 '같은 한계의 반복'으로 규정하며, 오히려 감독이 줄어든 만큼 코드 품질이 저하될 수 있다고 경고합니다. 핵심 문제는 이 도구들이 소프트웨어가 할 수 있는 일의 범위 자체를 넓히지 못한다는 것입니다. Jev는 자연어 의도를 상태 머신의 결정값으로 변환하는 새로운 원시 타입으로, 기존 로직 게이트(if/else/loop)에 확률적 판단 레이어를 추가합니다. 이는 소프트웨어 엔지니어링 자동화가 아닌 소프트웨어 능력 자체의 확장입니다.
AI는 이미 충분히 똑똑하다 — 자동화가 없는 이유는 다른 데 있습니다
Almeida가 제시하는 가장 날카로운 역설은 다음과 같습니다. GPQA(대학원 수준 질문 답변) 벤치마크는 이미 '해결됐다'고 불리지만, 패스트푸드 드라이브스루는 여전히 자동화되지 않았습니다. OpenAI는 2020년부터 고객 서비스 자동화를 시도했지만 2026년 현재도 실패 중입니다. 그의 진단은 RLHF 이후 AI 산업이 인간 평가자 최적화 방향으로 흘러가며 실질적 자동화와 멀어졌다는 것입니다. 즉, 신뢰할 수 있는 배포 가능한 AI가 아닌, 인간에게 인상적으로 보이는 AI를 만들어온 것입니다. 이 구조적 왜곡이 Jev의 설계 철학인 'prod, not God'의 배경입니다.
신뢰성은 업타임 SLA가 아닌 — 개발자가 예시 없이 믿고 쓸 수 있는 지능적 일관성입니다
Almeida가 말하는 신뢰성은 세 층위로 구분됩니다. 첫째, 서버 업타임 수준의 SLA. 둘째, 동일 입력에 동일 출력을 보장하는 결정론(determinism) — 그러나 UUID 하나만 바뀌어도 동작이 달라지는 LLM에서는 달성하기 어렵습니다. 셋째, 그가 목표로 삼는 수준: 개발자가 예시 쿼리 없이도 Jev를 신뢰하고 코드를 작성할 수 있는 '지능적 일관성'입니다. 이 세 번째 신뢰성이 실현되면 개발자는 'permaflow' 상태에서 AI를 인프라처럼 사용하게 됩니다. Almeida는 이를 위해 더 일찍 출시할 수 있었음에도 의도적으로 지연했다고 밝혔으며, 이는 데모 최적화가 아닌 실제 배포 가능성을 최우선으로 한 선택입니다.
SaaS는 AI에 의해 파괴되지 않는다 — 오히려 가장 큰 수혜자가 될 수 있습니다
코딩 에이전트 등장 당시 'SaaSpocalypse' 서사가 SaaS 기업 가치를 폭락시켰습니다. 그러나 Jev 출시 이후 SaaS 기업들의 반응은 정반대였습니다. Almeida의 해석은 명확합니다. SaaS 기업의 진정한 자산은 코드베이스가 아니라 수백만 기존 사용자에 대한 도달 능력(distribution)입니다. Jev 같은 지능형 레이어가 기존 소프트웨어에 내장되면, 멀티초이스 폼 같은 인터페이스는 사라지고 '내가 의도한 대로 해주는(do what I mean)' 경험이 됩니다. 새로 구축하는 경쟁자보다 이미 고객에게 도달한 기업이 이 전환을 더 유리하게 활용할 수 있습니다.
확률적 프로그래밍의 부활 — AI는 소프트웨어 아키텍처의 새 원시 타입이 됩니다
1970년대에 사실상 소멸된 확률적 프로그래밍(probabilistic programming)이 Jev를 통해 실용적 형태로 부활하고 있습니다. Almeida는 인터넷, 클라이언트-서버 구조처럼 새로운 프리미티브가 등장할 때마다 시스템 전체를 재건하는 기회가 생겼음을 상기시킵니다. Jev의 경우, 지능이 고수준 인터페이스가 아닌 시스템 내부(TCP/UDP 수준)에 내장되는 방향을 지향합니다. 그의 핵심 지표인 '달러당 지능(intelligence per dollar)'은 이 미래를 향한 설계 원칙으로, 인간 소비용 LLM 호출이 아닌 소프트웨어 내부의 수많은 자동화 판단을 위한 것입니다. 이는 단순한 제품 출시가 아닌 소프트웨어 아키텍처 패러다임의 전환을 주장하는 것입니다.
Coding agents are faster shovels — they don't expand what software can actually do
Cursor, Codex, and Claude Code produce the same kind of code a human would have written ten years ago, just more quickly. Almeida's critique is not about speed but about expressive power: these tools leave the fundamental capabilities of software unchanged. Jev is a new primitive — a library that lets developers embed natural language intent directly into program logic, returning a decision with confidence levels rather than raw text. This is the difference between automating the act of writing software and expanding what software itself can reason about. The analogy he uses is adding a 'small brain' to the existing three logic gates of programming.
The automation gap isn't a data problem — it's an incentive misalignment baked in after RLHF
The most uncomfortable argument in this conversation is structural. After RLHF demonstrated remarkable generalization in late 2021, the AI industry optimized for human evaluators rather than autonomous task completion. The result: models that score at doctoral level on GPQA benchmarks yet cannot handle a drive-through order, and OpenAI's customer service automation effort has been running since 2020 without delivering. Almeida's diagnosis is that 'over-promise, under-deliver' became the industry's default mode because human impressiveness and genuine reliability are different optimization targets — and the industry chose the former. Jev's design philosophy, including the deliberate decision to delay launch for reliability, is a direct rebuttal.
SaaS incumbents are AI's biggest winners, not its casualties
When coding agents emerged, 'SaaSapocalypse' became a credible narrative and valuations dropped. When Jev launched, the same companies celebrated — and the reversal is theoretically coherent. A SaaS company's durable asset is distribution: it already reaches millions of users. If intelligence becomes an embeddable primitive rather than a standalone product, incumbents can make their existing software dramatically more capable without rebuilding from scratch. Almeida envisions multi-choice forms disappearing entirely, replaced by software that interprets natural language intent directly. The startup that has to acquire those users faces a structurally harder path.
The real benchmark for AI progress is boring task automation, not benchmark scores
Almeida proposes a deliberately unglamorous test for AI maturity: can it reliably automate things that obviously should be automatable? His 'canary in the coal mine' is not whether a model can solve olympiad math but whether it can run a customer service queue without human oversight. He draws a sharp distinction between RSI (recursive self-improvement) — which he does not believe is on the near-term path — and automating 'the rote and simple work that basic instructions can handle,' which he argues models have been capable of for years. The bottleneck is not intelligence; it is the absence of a reliable interface between that intelligence and software systems.
Probabilistic programming is being revived — and will reshape systems architecture from the inside out
Probabilistic programming, a field that effectively died in the 1970s, is re-emerging through products like Jev as a practical engineering discipline rather than an academic one. Almeida frames this as analogous to the architectural shifts brought by the internet and client-server computing — moments when a new primitive forced a full rebuild of systems. His stated ambition is for Jev to operate deep in the stack, at the TCP/UDP level of abstraction rather than the UI layer, measured by 'intelligence per dollar' as the guiding metric. If this framing proves correct, the most consequential AI infrastructure may end up invisible to end users — embedded in the decision logic of software they already use.
자동화는 다 어디 있는 건가요? AI는 믿을 수 없을 만큼 똑똑한데, 정작 다른 모든 것에는 이토록 쓸모가 없습니다.
— Diogo Almeida, Founder & CEO, TypeSafe AI
Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff
— Diogo Almeida, Founder & CEO, TypeSafe AI