TypeSafe Founder: "Build Products, Not Gods" 鈥?Chasing "Extreme Reliability," Jev Set to Reverse the "Software Extinction Theory"

Deep News
09/29

Where exactly has automation gone? AI has become incredibly intelligent, yet it remains useless for everything else 鈥?and that is truly a tragedy.

Recently, TypeSafe founder Diogo Almeida stated bluntly on the a16z podcast that AI coding tools merely make code get written faster, but the software itself has not improved. OpenAI has been attempting to automate customer service since 2020 without success, and the root cause is that AI and software are "going their separate ways."

In a conversation with a16z co-founder Ben Horowitz and partner Martin Casado, Diogo laid out his assessment of the current AI predicament and what his company's explosively popular AI decision-making model 鈥?Jev 鈥?is trying to solve.

He explained that Jev 鈥?a new programming primitive that embeds intelligence inside software 鈥?enables programs to understand intent and make probabilistic decisions. Diogo believes extreme reliability is the prerequisite for everything, regards "extreme reliability" as the core of Jev, and advocates "building products, not gods."

Diogo judges that SaaS companies will be among the biggest beneficiaries of the AI wave, and the "software extinction theory" will be reversed. Jev's ultimate goal is to make all technology "do what I mean."

Diogo Almeida

Where Is the Automation: The Gap Between AI Capability and Real-World Deployment

Diogo believes that existing AI coding tools 鈥?whether Claude Code, Codex, or Cursor 鈥?produce code that is essentially no different from what was written 10 years ago.

No matter how many AI coding tools you use, the software itself has not actually gotten better. Perhaps you write faster, but arguably it has gotten worse 鈥?because there is less oversight.

A more specific example: OpenAI has been trying to automate customer service since 2020 and still has not achieved it. "That is crazy. We have such strong financial incentives to automate these things, but we just have not done it."

In his view, the root cause is neither data nor insufficient model capability, but rather that AI and software have always "gone their separate ways" 鈥?AI outputs natural language, software processes structured instructions, and the two have never truly connected.

What Is Jev: An Intelligence Primitive Embedded Inside Software

Diogo positions Jev as an entirely new software primitive 鈥?an intelligence layer embedded within code, rather than a tool to replace software engineers.

What I want is not to automate software engineering, but to expand what software itself can do, so that things that should be automatable can truly be automated.

Jev's approach is to provide a new programming primitive 鈥?developers can describe intent in natural language, give it a state machine, and it will decide what to do with a certain level of confidence.

It is like a library where you can describe requirements in natural language, hand it a state machine, and it will choose what to do with confidence 鈥?a capability that has never been this widely available before.

He used an analogy to explain the difference from coding agents: coding agents are good at syntax, poor at semantics, and terrible at architecture. Jev solves a problem at a different level 鈥?allowing software to understand intent internally, rather than merely generating text for humans to interpret.

Diogo also candidly acknowledged Jev's essence: "Jev is absolutely a classifier. Classifiers are designed to be useful in the first place." He believes Jev's capability may already surpass what a dedicated MLE team in 2019 could have built as a narrow model for a specific task.

Extreme Reliability Is the Prerequisite for Everything

Diogo repeatedly emphasized that Jev's core competitiveness is not performance, but reliability.

Reliability is everything about this thing. If you do not understand this, it is very hard to make a truly copyable competitor 鈥?it is not something you can replicate just by running benchmarks.

His definition of reliability has three layers:

Availability (uptime/SLA): This is the foundation, similar to an SLA 鈥?whether the system can run at all.

Robustness: Every call demonstrates a similar level of intelligence 鈥?not strict determinism, but "smart enough every time."

Trustworthiness: "Do what I mean" 鈥?developers can rely on it without repeatedly providing example queries, trusting it directly.

The highest honor of reliability is reaching a state where developers do not need to provide example queries when using Jev 鈥?when you trust it directly, you enter a flow state.

He also stated bluntly: "We could have released much earlier. I do not think people realize this. But it brings that feeling of 'I can trust this thing' 鈥?it is an anti-frustration machine."

"Build Products, Not Gods": Divergence from the Mainstream AI Narrative

Ben Horowitz pointed out the most striking difference between TypeSafe and other AI companies: "My favorite thing you say is 鈥?we build products, not gods."

Diogo responded directly: "I do not think we are on the path to recursive self-improvement (RSI). I do not think so now, and I never thought so before."

He attributes the current pessimistic narrative in the AI industry to the "superstition of a single model" 鈥?the belief that one super-brain can rule everything. "That is what people hear, but it is not reality."

His judgment is that the AGI defined by OpenAI 鈥?"automating most economically valuable work" 鈥?is technically achievable, but the industry has gone in the wrong direction.

Since RLHF, the AI industry has split into enormous over-promising and under-delivering. Because humans are evaluating how good models are, the models look great, but we have been optimizing this evaluator rather than automation itself.

His goal is more modest:

What I want is "do what I mean." Imagine if all technology could do what you mean 鈥?that is not science fiction. Just look at how smart AI is.

SaaS "Reverse Doomsday": The Software Extinction Theory Will Be Reversed

When coding agents emerged, the market once feared SaaS companies would be disrupted, and valuations fell sharply. After Jev appeared, the reaction from SaaS companies was the opposite 鈥?they generally welcomed it.

Diogo's explanation is that the premise of the SaaS extinction theory 鈥?"software is cheap and easy to replicate" 鈥?never really held.

A lot happens under the hood, and the value software provides has not disappeared.

He believes SaaS companies will be the biggest winners of AI capability deployment, for three reasons:

They understand user workflows best and know what is worth automating.

They have already reached a large number of users, and the distribution cost has already been paid.

They have the ability to embed Jev into existing products, truly enhancing software capability rather than just adding a chat box.

I think this will be a reverse doomsday. I am very excited about it.

He gave a specific vision: multiple-choice forms will disappear, replaced by software that directly understands user intent. "'Do what I mean' will be pushed to an entirely new level."

Full Interview Transcript:

How Jev Turns AI into Software That Actually Gets Things Done

a16z podcast

a16z's Ben Horowitz and Martin Casado sat down with TypeSafe AI founder Diogo Almeida to explore a simple question: AI has become extremely powerful, so where exactly is the automation?

Diogo believes coding agents may help us write software faster, but the software they produce is essentially no different from before. TypeSafe is taking a different path with Jev: embedding intelligence into software itself so developers can build programs that understand intent and make probabilistic decisions, rather than merely generating text for people to interpret.

They discussed why reliability is the key to making AI truly programmable, how this will open a new era of probabilistic software, and why mature SaaS companies may be particularly well-positioned. Ultimately, Diogo's goal is simple: to build technology that can reliably "do what I mean."

Chapter 1: Opening

Ben: Where exactly is the automation? AI is incredibly smart, yet so useless for everything else.

No matter how many AI coding agents you use, the software itself has not actually gotten better. Perhaps you write faster, but arguably it has gotten worse.

Diogo: OpenAI started trying to automate customer service in 2020. What I really want is intelligent software 鈥?I want to expand the boundaries of what software itself can do, so that things that should be automatable truly become automated.

Ben: My favorite thing you say is: "We build products, not gods." That is great. Because any other leader of a major lab, even with genuine enthusiasm, would wrap everything up. But your perspective is completely different.

Diogo: Exactly. We want to create a better world, for many nuanced reasons. I do not think we are on the path to recursive self-improvement (RSI), and I do not buy that "SaaS doomsday" story either.

Chapter 2: Getting to Know Diogo and TypeSafe

Ben: Today, TypeSafe founder and leader Diogo joins us. He is a heroic figure to both Martin and me. He is not only building a genuinely interesting product, but also driving a movement we believe is profoundly important. So we are very excited today. Welcome, and thank you for coming on.

Diogo: Thank you.

Ben: Maybe you can give us a brief introduction 鈥?what is Jev? What is TypeSafe? Why does it matter?

Diogo: Okay, let me explain. Someone actually asked me to do an elevator pitch, but I usually ramble and do poorly. But I realized my favorite elevator pitch for Jev is: Where exactly is the automation?

It is truly deeply regrettable. So much wisdom, AI is so incredibly smart, yet so useless for everything else 鈥?it is painful. I am not attacking chatbots or coding agents; I love them myself. But for everything else, they just cannot help, and that is sad.

We have so much raw potential that has not been polished into something useful for work 鈥?this is exactly what TypeSafe is doing: building AI for software. We want to make AI not just powerful for the humans in the loop, but truly capable of building practical, usable software. Jev is our first model in this space, with the goal of making automation better, better, better.

Ben: Yes. What makes this interesting is that it caused a sensation in the software world. Every developer we know called and said: "This is amazing, fantastic, fast and good, improved in every way." But people also ask: do we not already have Claude Code, Codex, and other tools? What is the difference? How does this lead to real automation?

Diogo: I wish I had a nice diagram handy, because I have one I really like to explain this. I appreciate Claude Code and Codex, and I like Gary Tan's description of them 鈥?"instant software," which is a brilliant way to describe what they do. They generate software instantly, letting you program in natural language with expressive power comparable to traditional software.

But what I want is something different 鈥?intelligent software. What I want is not to automate software engineering, but to expand what software itself can do, so that things that should be automatable truly become automated. In more vivid terms, I want to express the concept of "intent." I want to expand the vocabulary of what we can express and achieve.

The essence of programming is hyper-precise specification of valuable things, then infinite replication. That is remarkable. I just want to make all of this more and better.

Chapter 3: Intelligent Software, Not Faster Code

Ben: Oh, interesting. So one way to understand it is: you are not talking about replacing software engineers with a faster, possibly inferior tool, but rather 鈥?no, no, we want to give existing software engineers a super engine so they can write better and more interesting things.

Diogo: Yes, exactly. Actually, I think many people miss this point. It is so subtle yet so important, and worth elaborating on.

If you use tools like Claude Code, Codex, or Cursor, they can write code, but the code they produce is essentially no different from what humans write 鈥?perhaps better, perhaps worse, but fundamentally the same kind of code from 10 years ago.

What Jev means is that whether you use Claude Code or write by hand, you gain a new foundational primitive 鈥?a new thing that can be embedded in your code to truly expand software's capabilities. So it is not generating code, but being contained within code...

Martin: Yes, and this is a very powerful foundational primitive 鈥?but it is also somewhat different from how programmers usually think. For example, it introduces the concept of probability, an intelligence layer embedded inside software.

Diogo: Yes. Think of it this way: it is somewhat like a library where you describe what you want in natural language, give it a state machine, and it will choose what to do with a certain level of confidence 鈥?a way of doing things that was not as widely available before.

There is a lot to dig into here. I want to say one thing first: I really love the direction of the question "where exactly is the automation." I love writing software; I could write all day. Of course, I would not recommend anyone become a CEO, regardless.

And this is indeed crazy 鈥?AI is so powerful, yet software has barely changed in 10 years, and nobody can reconcile these two things. The most we can do is add a chatbot on the side that can execute some operations, but not all, because some operations are still not reliable.

I want to return to that point about "different ways of thinking." Yes, I think machine-native things and code cannot fully match precisely 鈥?and that is exactly the art we are working on. On a new employee's first day, I draw a Venn diagram: what AI is good at, what is valuable in code, and we are in the intersection.

So we do not output things like inferring floating-point numbers, because AI is genuinely bad at that. But things like probability are not new, and this leads to the argument that "Jev is just a classifier."

Jev is absolutely a classifier. Classifiers are powerful; they are designed to be practical. They are designed to make systems work, and the interface is the same as some machine learning concepts 鈥?because these all come from practitioners trying to make systems actually work.

And what I am seeing now is that Jev actually 鈥?my guess is 鈥?may already be stronger than assembling a dedicated MLE team in 2019 to do these things for you, and you can program and configure it in real time. Who knows what else can be built, because in 2019 there were not many excellent MLE teams that could build narrow-task systems. Collecting datasets, evaluating results 鈥?all of it was hard. And this is just the beginning.

Martin: I feel there is a lot of possibility here. By the way, do you think there is a slider here? One end is "language in, language out," like what we have today; the other end is "existing imperative programs," and you can move between the two. Or do you think "language in, state machine out" will solidify as a general tool in the design space?

Diogo: That is a tricky question. Let me give you my honest answer 鈥?my honest answer is that it really is a slider.

When designing various properties for the product, I may have made some mistakes due to personal preference. But currently, "intelligence per dollar" is my north star metric 鈥?of course, that may not be right, and "intelligence per second" may be more valuable in the short term.

But even our interface design 鈥?calling the input "state" 鈥?is intentional: it is part of the program's internal state. Deep down, we are actually doing a lot of optimization work for more complex arrangements of program internal states.

Can we put intelligence in? I think this will be an endless exploration. We are very careful in our design, and from a practical standpoint, certain things will naturally happen 鈥?for example, making AI decisions at the millisecond level is easier to implement, so for a while it will be more like a database than a standard library thing. But I would also love for it to become part of the standard library someday.

Chapter 4: Diogo's Background and Views on AI

Ben: Can we step back and talk about what made Diogo? You speak like an AI researcher, a systems engineer, and a programmer 鈥?groups that usually have very little overlap. You have turned AI 鈥?something we have been pushing toward "human-like existence" 鈥?into a tool for programmers. Can you talk about your personal journey?

Diogo: My path into AI is somewhat unusual. I was a math competition competitor and won awards. How to describe it... My math was good enough to attract girls, which is already quite impressive.

Ben: Is that really a thing?

Diogo: Absolutely true. So you have to practice to a fairly high level.

Martin: So with math that good, what kind of girls can you attract? Our audience needs to know this.

Diogo: No, no, we need to inspire young people.

Martin: Young people, do not learn this. Just stay handsome, relaxed, and interesting. Do not try to compensate with talent.

Diogo: Ha, I cannot believe I said that. Anyway, I was a math competition competitor, but somewhat shamefully, I never actually liked math and never worked hard at it 鈥?just a small fish in a big pond. Math was always just a prescribed competitive path for me. I hated it because it was only about winning competitions.

Later I discovered that computer science is actually a lot like math, but cooler, more useful, and more interesting. I still love creating algorithm interview questions; that is one of the most beautiful things for me. It lets me enjoy myself and also very effectively assess others' abilities. So I love computer science, and I consider myself first and foremost a computer scientist, not an AI researcher.

I really entered this field because I won a Kaggle competition 鈥?not through complex math, but by pushing automation to the extreme: more nested loops, more... like solving it with a systems approach.

That eventually forced me to give a talk at NeurIPS 鈥?usually an honor, but I hated it because I just wanted to keep my head down and work.

Ben: Was that the Kaggle competition?

Diogo: Yes. The Kaggle competition was organized by Isabel Guyon, co-inventor of the SVM 鈥?should be first author, but I am not 100% sure. She saw that I was the kind of person who did not quite fit into academia, took me in, introduced me to various people in the AI world, and my career trajectory was pushed in this direction.

Martin: And then OpenAI?

Diogo: No, first a startup founded by Jeremy Howard.

Martin: Really?

Diogo: Yes, I really like Jeremy. Then Google Brain for a while. Then I retired for a bit, and eventually found doing nothing boring, so I thought: AI is actually really interesting, so I joined OpenAI. It turned out very well, really very well.

Ben: That is remarkable. You just said something extremely rare in today's world: "AI is really interesting." And your company's spirit and view of AI are so different from others. My favorite thing you say is: "We build products, not gods." That is great. Any other leader of a major lab, even with genuine enthusiasm, would wrap everything up. But your perspective is completely different.

You say: no, we want to create a better world, and it will be great 鈥?not fewer jobs, but more and better jobs, and everyone will have a bright future 鈥?and being with you, you can feel that you genuinely believe this.

So please tell us, TypeSafe and Jev are not just a company to us, but a movement toward a positive future 鈥?and that is exactly the direction most people in the AI world do not buy into.

Diogo: Yes, or they simply do not understand it. That complaint that "Jev is just a classifier" is entirely an ML-level concern 鈥?while everyone else is celebrating, because everyone is thinking: this is amazing, we can finally do what we have always wanted to do!

If you do not truly understand developers, it is hard to understand what is happening. I completely agree. I do think there is a fairly negative world picture circulating, and I obviously do not agree with it.

I think this fundamentally comes from a superstition of a "single super-model" 鈥?everyone believes one all-powerful brain can rule everything.

Martin: That sounds even more ominous. That is the fear in people's minds.

Diogo: Exactly, that is how people understand it. But is that "one brain" really on the path to ruling us all? We have not even automated many basic things 鈥?things that we should not be having humans do.

There is a lot of very basic work that has not been automated. And every time reality diverges from this narrative, I feel great pain. Part of this disconnect comes from the question "where exactly is the automation" 鈥?AI is clearly so smart, and the economic incentives for automation are so strong, yet reality lags so far behind.

This gap between reality and AI's potential is what I truly lament, and it is the driving force behind my desire to do something. Now, we have finally released Jev, and for me it feels like the start of a party. Developers using Jev are the happy AI crowd, while those not using it are gloomy 鈥?it really is a fascinating contrast.

Ben: This reminds me of a conversation this morning with David George, who leads our growth fund. We were talking about new tools, and I asked if he had tried Perplexity's Muse. He said it was great. I asked what he used it for, and he said: "I finally canceled my New York Times subscription." I said that is actually pretty hard to do. But this is just the tip of the iceberg 鈥?there is a huge amount of annoying stuff waiting to be automated.

Chapter 5: Is This Just a Data Problem?

Diogo: I think if we are truly going to be honest about the north star goal of automation, we cannot repeat AI's long-standing mistake 鈥?focusing too much on outliers and demos.

Many people ask me what my favorite use case is, but I am not even sure those use cases really work. I want it to run quietly in the background, so people can trust it without having to watch it constantly, and others can build on top of it.

At the same time there is security 鈥?another kind of safety: if you want it to actually get resources, gain access, and execute tasks, you need guarantees, or at least statistical guarantees, that it will not go out of control.

Of course, I do not think our model will go out of control anytime soon, unless someone specifically writes software to make that happen 鈥?that would be a very cool feat. But regardless, that is not the direction we should be responsible for.

Martin: I am curious, how long has this intuition been brewing? I remember we talked about it around 2017...

Diogo: Yes, we did talk about it. At that time I talked about the importance of data and the idea of focusing on tasks. I just want to ask, did all of this already point toward classifiers, or was it just an intuition 鈥?that there is another way to look at the entire AI movement?

Interestingly, that 2017 talk actually had a similar theme, something like "modular in theory, flexible in practice" 鈥?very software-style. So I have been fairly consistent on this.

I think it really started shortly before ChatGPT was released. When we released those models, I did not have much of a premonition about it. Honestly, I was even very pleasantly surprised by the generalization ability of RLHF.

Ben: When was that?

Diogo: Around late 2021, fourth quarter of 2021. That generalization ability was really strong 鈥?if you read that paper, it was different from other papers trying to prove their own points. We truly tried to falsify it with something close to the scientific method 鈥?for example, is this some kind of cheating?

One of my favorite test queries was: "Why is it important to eat socks before meditating?" We confirmed in advance that this sentence had not appeared on the internet, and the model actually gave a plausible, human-like answer. That triggered the realization in the team: this is not cheating.

In ML, you always have to be careful about the possibility of cheating. Then what really frustrated me was that we released that model 鈥?I am a firm capability accelerator, I did a lot of work to release that model 鈥?and I genuinely thought there was a considerable probability that model was AGI. But when it did not reach that level, my entire worldview collapsed, and I asked myself: why?

Ben: So at that time you were actually on the "crazy train" to some extent?

Diogo: Not exactly. I just thought RLHF generalized quite well, and maybe we already had AGI. RLVR is the thing that, in my current view, does not generalize as well.

Speaking of the definition of AGI, in OpenAI's early days 鈥?around 2020 鈥?people would describe AGI as "Ilya (Sutskever) and every if statement." But that was just an intentionally vague concept, so everyone could stand under this big tent and work together.

And I 鈥?for nuanced reasons 鈥?do not think we are on the path to RSI (recursive self-improvement). I still do not think so.

But I do think the AGI defined by OpenAI is entirely achievable 鈥?"automating most economically valuable work in the world." That actually sounds... there is a lot of work out there, much of it very mechanical, simple, repetitive work. To outsource work, you need simple, understandable instructions, and that level of intelligence has long been present in models, for quite some time now.

And what gnaws at me is: why has this not been put to use?

Since RLHF, the AI industry has essentially moved toward a mode of enormous over-promising and under-delivering. GPT-3 was actually fairly accurate at the time, but because humans are the judges of model quality, and the model performed well in human evaluations, it seemed like everything was great 鈥?but what we have been optimizing is the "judge" itself, not automation capability. That is the missing piece.

So it was from that point that it truly hit me 鈥?why is this thing not more useful?

Martin: So you think the metric should be: how much can you automate truly productive tasks? That is the dimension you mean by "over-promising and under-delivering"?

Diogo: In my mind, it is about whether the cool science fiction vision can come true. I think the ability to automate real tasks is the "canary in the coal mine" for that science fiction vision.

Are you really telling me math has been solved? Even GPQA (Google's anti-cheating Q&A) was solved two years ago, yet we still cannot handle drive-through ordering? These two things are hard to hold together in your mind at the same time, and I think many people do not have a good answer for it.

Chapter 6: The Breadth of Use Cases

Martin: Can I test an idea? Maybe it is not quite right, but 鈥?is the real-world distribution different from the digital world? The real world is heavy-tailed, with many exceptions, and we do not have the corresponding data. Could it be precisely because we do not have data from that distribution, and have not trained on that distribution, that AI is basically confined to low-dimensional manifolds like math or code?

Diogo: I do not fully buy the data argument. I believe there is indeed a long tail, and denying that would be absurd. But I think in the "canary in the coal mine" scenario, we do not need to automate that long tail.

I think we need to be extremely pragmatic about everything. Building reliable software is always an investment, like... what are the three virtues of a programmer? Laziness 鈥?unwillingness to repeat the same thing; arrogance; and the third one...

Martin: Yes, yes, I remember this is from the Perl era... and the third one.

Diogo: Right, I cannot remember either. But it is about that kind of laziness 鈥?being willing to spend 10 hours solving a 5-minute task so it never has to be done manually again. That should be an ROI decision for people who want to automate.

I just want to make automation possible, and I believe people will create new kinds of work from it 鈥?like a "Jev in jeans" kind of role. But as a benchmark, I think it is valuable to verify whether things that AI appears able to automate can actually be automated.

OpenAI has been trying to automate customer service since 2020 and still has not succeeded 鈥?that is truly shocking and baffling.

Ben: And within companies, very few things are actually automated now. Most projects have not succeeded, except programming, which has done very well.

Can you help us categorize what types of problems are easier to automate now? This is interesting because we studied customer service before this generative AI wave. You visit a company, they say "we solved 95% of help desk calls," which sounds like a lot, but when you look closely at the data, it is basically all password resets. If you break it down by uniqueness of the problem, maybe only about 50%.

So when you are dealing with humans and natural systems, there is this endless long tail of exceptions.

Diogo: Yes. Actually, about every hour someone comes to me and says "I am using Jev for this new use case," and it blows my mind. So, you asked whether I anticipated such a broad range of use cases 鈥?I did not expect this. This release was something no one could have predicted. Who could have predicted it? It is a bit crazy.

It is like none of us could have predicted ChatGPT, the consumer-facing thing, and this is the developer-facing version, which feels even stranger. I do not even know what proportion of people at the "Jev party" are developers. I cannot imagine how non-developers use it, but even my non-developer friends are celebrating on Twitter, posting memes, fully immersed.

Okay, first point 鈥?amazing. Second point 鈥?this is hard to convey in a short sentence, because behind it are years of blood, sweat, and tears. My obsession with reliability is very, very high. Reliability is the soul of this product. If you do not understand this, it is very hard to make an imitation that looks similar on benchmarks.

I think every "nine" of reliability will be incredibly valuable to everyone 鈥?even if it is not the most valuable thing in market cap terms, because it will unlock new applications. We are fighting for various exotic orders of reliability, some of which we ourselves do not fully understand yet, because we are like installing AI's intelligence engine into people's workstations and letting them discover what it can do.

Martin: What does reliability specifically mean here? Is it model availability, returning the same result on a call, or something else?

Diogo: Good question. For something that is inherently stochastic 鈥?

The first is availability/SLA, which is not quite this. The second is closer to determinism. The third, I would call robustness 鈥?similar intelligent performance every time.

Not exactly determinism, because determinism is useful for unit tests but not enough for real systems. For example, if you add a UUID to the prompt, it is functionally the same, but the result will not be exactly the same. There is another layer I do not yet know what to call 鈥?maybe some kind of intelligence: it does not need to perform exactly the same operation every time, but it needs to be smart every time. If you were in that situation, is this a reasonable judgment a human could understand? Developers can program on top of that.

For me, the highest honor of reliability will be reaching a state where developers can program directly against Jev without writing any example queries 鈥?because they fully trust it and enter that "flow state."

Martin: This makes me think that with such a foundational primitive, the value of coding agents actually decreases somewhat 鈥?because even if Codex writes all your software for you, if that software does not use Jev, the software's capability is inherently limited. Conversely, even if you do not use coding agents, as long as you use this general foundational primitive, writing software becomes easier.

Do you think the future is coding agents using Jev, or more human-led?

Diogo: This is more a question about coding agents than about Jev. My feeling is that I have not used them as deeply as I would like. You two probably use them more than I do 鈥?which I regret a bit.

But my experience is that they are very good at syntax, quite poor at semantics, and terrible at architecture.

Martin: Yes, very poor at architecture.

Diogo: And architecture, to me, is the most human-creative part of software. So I really like using coding agents. I think Jev is probably not yet in their training distribution 鈥?if they actually trained on our user data, that would be scary, but they probably did not.

But once Jev enters their training distribution, I think there is no problem letting them handle syntax. As for architecture, maybe the model is not useless, but at the 50th percentile. If you do not understand architecture anyway, maybe that is good enough. These are all gray-area trade-offs. Sometimes speed is the knob your company or project needs most 鈥?you are willing to accept 50th percentile architecture instead of 60th percentile in exchange for faster speed, letting Codex run overnight.

Chapter 7: The SaaS Doomsday, Reversed

Ben: Speaking of this, an interesting phenomenon has already appeared in the market: when coding agents emerged, people shouted "the SaaS doomsday is here," and all SaaS company valuations fell through the floor. But when Jev appeared, every SaaS company said "this is the best thing ever." How do you explain this reversal?

Diogo: I think it is very natural, not much to say.

In the SaaS doomsday narrative, I think the argument that really does not hold up is: software is very cheap, and perhaps easy to replicate. I can believe the former to some extent, but I completely do not believe the latter, because a lot of value happens under the hood. Maybe I am a bit too enamored with software here.

Martin: We all are.

Diogo: Okay, okay. So I do not think that narrative holds. The value SaaS provides seems the same as before; maybe the market was just scared.

I think SaaS will be one of the biggest winners in the entire AI game. I really want to work closely with the largest, most "boring," most user-pain-aware SaaS companies, because they know best which workflows should be automated and what users truly need 鈥?that is their core competency.

Software is always a capital investment, but you spend money upfront to make the experience better, then distribute that improvement to all those massive user groups. So I think, from a capability dimension, this will be a "reverse doomsday." I am so excited, I need to give it a name.

Ben: Yes, it needs a name. "SaaS Carnival"?

Diogo: That sounds a bit too cheerful.

Ben: All SaaS applications will suddenly become dramatically more useful. And a lot of a SaaS company's capital investment is really about reaching all customers. If you have already reached all customers, and then instead of just adding a chatbot to the product, you make the software itself better and better 鈥?that would be remarkable.

I do not know if this is a realistic dream, but I think there is a possible world where those multi-select forms disappear. I feel like they are just mapping natural language to structured output that software has long been able to handle 鈥?just in a different wrapper.

Diogo: And strictly speaking, this concept comes from the 1980s. Back then we called it "fourth-generation language" (4GL). Do you remember?

I think "Do What I Mean" will be pushed to an entirely new level. If I can give an example of a Jev application 鈥?I am not sure it is reliable enough, so I dare not promise 鈥?but someone used voice to control a computer, and the system continuously judged in real time: is this a command, or text to insert? What text is being inserted? That sounds extremely cool. Interfaces may completely change, though of course we still need to make it cheaper and faster.

Ben: That would be Star Trek level.

Martin: There is a profound intuition here: if you use AI to generate software today, you are still creating the same software as before. But look at an ordinary PR in a big company 鈥?on average about 10 lines of code 鈥?we actually did this research. So you automated 10 lines of code, and those 10 lines may just reflect some customer learning.

What you are optimizing is actually a fairly small thing, and it does not give software new capabilities 鈥?you are just automating this inherently small thing, which in the limit has quite limited significance.

But now, with this new capability, software itself can truly get better.

Diogo: Yes. And honestly, before Jev, I never realized: no matter how many AI coding agents you use, the software itself has not actually gotten better. Perhaps you write faster, but arguably it has gotten worse, because there is less oversight.

Ben: And often less secure.

Diogo: Yes, definitely. But now, you can truly say: applications will gain new functionality because of this new foundational primitive you provide. It can understand natural language, reason, and combine that ability with state machines.

If people can take this as the biggest takeaway, that would be the greatest compliment to our work. I feel that extending the existing three logic gates to more is almost too grand a vision 鈥?in our type system, some types are like the same logic gate, but with a tiny brain inside. If this truly becomes TypeSafe's legacy to the world, it would be something extremely important for the world. I will not over-promise, but I will fight for it.

Chapter 8: Application Layer or System Foundation?

Martin: One question: in serious scenarios that require strong guarantees, like state consistency, durability, and true system-level concerns, how deep can this go?

Analyzing logs, analyzing emails, providing UI, talking with humans 鈥?these are definitely fine. But over time, will this become something like an intelligent database? Even... air traffic control systems 鈥?which we really need.

Diogo: That is indeed a bit scary. My philosophy has always been: automate the easy things first, then the hard things. But I also think an entirely new era of probabilistic programming will open up. Probabilistic programming actually has a long history, and basically died out in the 1970s...

Martin: I am very familiar with this. Actually, you could also call Jev a kind of neuro-symbolic system. Your co-founder Eric comes from that background; he talked to me about Bayesian methods.

Diogo: Yes, he did a lot of work in biology, with ups and downs. But I mean something broader. My brand is pragmatism, extreme pragmatism. I do not like biologically inspired things very much. I think they have never truly worked 鈥?they are useful, they can inspire crazy people to work for decades until one day there is a real breakthrough, which then gets engineered and polished. The AI neural network story is indeed like that, but many explanations of "why it works" are actually inaccurate. For example, hierarchical feature extraction did work, otherwise residual connections would not be useful 鈥?but that is another topic.

From a systems perspective, I am not personally excited about this part 鈥?not because it is bad, but because it is just too complex to program. But I am excited for the world: when we have many different layers of intelligence across different costs and speeds, the true systems maniacs will make all kinds of crazy trade-offs at the extreme 鈥?Jev may be a thousand times too intelligent for them; they just need an approximate signal to make an approximate routing decision.

That will be extremely crazy stuff. And the good news is that we can rebuild systems again 鈥?we have a new foundational primitive, a completely new way of thinking about software. We did this with the internet, from mainframes to client-server architecture, and we do it periodically.

And, just because of cybersecurity issues, we will probably need to rebuild almost all systems, simply to make them secure.

Ben: This is especially clear for critical infrastructure.

Martin: When you think about Jev's applications, do you think more from the application layer, SaaS, analytics angle, or from the system foundation angle, or all of the above?

Diogo: I have a bit of my own view. My framework is somewhat like the foundation of TCP/IP 鈥?UDP is unreliable, TCP is too reliable, and in between...

As for why I focus on "intelligence per dollar" 鈥?I work backward from the science fiction vision of an AI-driven economic revolution, with AI everywhere and in all software. I ask myself: in such a world, of all AI calls, what proportion are "for human consumption" 鈥?needing style, tone, etc.? In the limit, that proportion is actually very small.

Starting from the same question: how many calls will be at the first layer, and how many will be deep in the foundation? I think the vast majority will be at the foundation, but it will start from the first layer. And if you do not aim for the foundation, it will be hard to get there.

I think many people do not realize how AI and software "pass each other at night." Even if you try to embed AI into software, its behavior will be strange 鈥?because software does not accept natural language input. You stuff the JSON format and schema you want into the prompt, and it just does not listen. In the end, you can only hand the output to a human, or to another LLM 鈥?that is the agent while loop.

So from first principles, there are only two paths: human in the loop (chat), or agent (while loop) 鈥?because natural language needs to be fed back into another model repeatedly.

I have seen too many people go through the "five stages of grief": pick up AI, intend to embed it into software, then denial, anger, and finally acceptance 鈥?forget it, I will just hand it to another LLM, hand it to a human.

And I think this is the first time I truly see: you can take an LLM, take AI, and truly map it to a state machine, and do so efficiently.

Of course, I do not want to over-promise. I do not know if it is ready for all those over-hyped use cases, but my team will fight for it. We really, really care about reliability.

Chapter 9: Reliability and "Do What I Mean"

Diogo: We could have released much earlier. I do not think people realize this, and I do not think they will realize it from discussions on Twitter. People will probably never truly understand, but it will have that feeling 鈥?oh, I can trust this. So it is an anti-frustration machine.

"Do what I mean" 鈥?for me, this is about the fluidity of the world: making everything run more smoothly, meshing tightly like gears. I actually have a whole "AI utopian vision" spread across different dimensions. "Do what I mean" is a very important part of it 鈥?imagine all technology operating according to your meaning. That is not science fiction. Just look at how smart AI is, right?

Martin: Truly amazing. Perhaps that is the best closing line for today.

Ben: "Do what I mean." I love it. Thank you, Diogo. This was a great conversation, really enjoyed it.

免責聲明:投資有風險,本文並非投資建議,以上內容不應被視為任何金融產品的購買或出售要約、建議或邀請,作者或其他用戶的任何相關討論、評論或帖子也不應被視為此類內容。本文僅供一般參考,不考慮您的個人投資目標、財務狀況或需求。TTM對信息的準確性和完整性不承擔任何責任或保證,投資者應自行研究並在投資前尋求專業建議。

熱議股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10