Database pioneer Mike Stonebraker criticizes industry and AI agent trends.

icon MarsBit
Share
AI summary iconSummary
Market trends reveal increasing scrutiny of database strategies, as Turing Award winner Mike Stonebraker criticizes Oracle, Google, and Amazon. He argues that AI agents encounter fundamental database challenges when transitioning to read-write operations, with large language models struggling to generate accurate SQL queries in real-world data environments. The Fear and Greed Index remains volatile as industry shifts fuel debate. Stonebraker’s remarks underscore ongoing tensions in the evolution of data infrastructure.

If I were to start over today, I’m not sure I would still recommend that an 18-year-old learn computer science.

The person who said this is Mike Stonebraker, a Turing Award winner in the field of databases, commonly translated into Chinese as "Shi Potian." He is one of the key creators behind Ingres and Postgres, and one of the most important figures in database technology. In his view, computer science may no longer be a growth industry in the future.

In this interview, Stonebraker called out nearly the entire database industry.

He criticized Oracle, directly accusing Larry Ellison of lying back then: selling customers features that weren’t yet implemented, portraying the future as the present, and then relying on early customers to help debug.

He criticized Google, calling its earlier promotion of MapReduce and eventual consistency "stupid." Many people blindly trusted Google simply because "Google is smart," assuming it always knew what it was doing. But in Stonebraker’s view, Hadoop was absurdly inefficient, and eventual consistency was only suitable for a very narrow set of use cases. When Spanner emerged, Google essentially admitted that fundamental database issues like transactions and consistency simply cannot be avoided.

He also criticized AWS: Amazon maintains around 15 different databases, but he believes only about three are truly necessary. Graph databases and various databases with redundant functionalities, in his view, lack sufficient performance or market justification to continue existing.

But more interesting is his perspective on today's AI wave.

In his view, what is currently called agentic AI is essentially “a large model plus a layer of system packaging,” and most are still stuck in a “read-only” phase. Once it enters the true “read-write” world—such as transferring funds or updating inventory—the issues immediately revert to classic database problems: transactions, consistency, and atomicity. This is not an AI problem, but a distributed database problem.

One more point is his judgment on using large models to write SQL.

On public benchmarks, the model achieves over 80% accuracy, making it seem just one step away from production. But in tests using real data warehouses, the result was 0%. Even with RAG added, or by directly feeding join conditions into the model, accuracy barely reached 35%. A skilled human engineer, by contrast, can achieve over 90%. He therefore concluded outright: this technology is not yet ready for production, at least not in the foreseeable future.

Below is the full interview.

1 Postgres: The best place to start, not the end

Host: I’d like to start with the origins of PostgreSQL, but before that, I’d rather begin at the very beginning: how did you get into the field of databases?

Mike Stonebraker: After graduating, I was fortunate to be hired by Berkeley. At the time, I was well aware that continuing in the direction of my Ph.D. research offered little future—then or now. The best path was to find a knowledgeable mentor who could guide me into the field.

So Gene Wong took me under his wing and said, “Let’s build something together.” It was 1971, the year after Edgar F. Codd published his groundbreaking paper in CACM, the Communications of the ACM.

Gene suggested we look into databases. At the time, there were primarily two camps: one was the Codasyl proposal, which you may not have heard of—it was a low-level, "spaghetti-like" network structure requiring pointer traversal to execute queries; the other was IBM’s solution, known as IMS, a hierarchical data structure essentially based on trees.

At the time, IBM also recognized that the tree structure was not universal and could not solve many problems, so they added some extensions to transform it into a restricted network structure. But this was clearly a very poor patch.

Codasyl also had many issues: it was very low-level, difficult to debug, and once your schema (though it wasn’t called that at the time) changed, you essentially had to start over from scratch, because it was entirely tied to the physical layer.

In contrast, Codd’s relational model was very sound. So Gene said, “Let’s implement this—that’s the next step.” We started Ingres in 1972, when I was still an assistant professor at Berkeley. You know, assistant professors typically have about five years to prove themselves—either earn tenure or be let go. Ingres was the key project that secured my tenure, which I received in 1976.

That’s how it all began. Later, other opportunities arose. At the time, many people built prototype systems—essentially code at a student-project level: it worked for them, but no one else could use it. We completed the first 90% and got it running; then, for reasons unknown, we spent an additional “90%” refining it into a truly usable system.

The Berkeley version of Ingres was actually usable. Over the next few years, about 100 universities began using it as Unix gained popularity. It was a free database system that ran on Unix and became very popular in academia. As a result, many people started visiting Berkeley, saying, "This is cool—what’s your biggest use case?" But we could only say that, in reality, it wasn’t very large.

This issue was fully exposed in a project at Arizona State University. They considered using Ingres to manage data for 40,000 students. They were willing to accept Bell Labs’ unofficial operating system or our “unofficial” database system, but the project ultimately failed because there was no COBOL available on Unix, and they were a COBOL shop.

Unsupported operating systems, unsupported database systems, and no COBOL—this directly renders us irrelevant.

The only way forward was to start a business. So in 1980, we secured venture capital at the time and founded Ingres Corporation, migrating the system to “real operating systems” like VMS and offering commercial support—marking the beginning of its commercialization.

Host: I see that Ingres was competing with Oracle Corporation at the time. Technically, you were clearly superior, but Oracle still won—how did they manage that?

Mike Stonebraker: Larry Ellison is an incredibly skilled salesperson. He makes no distinction between "now" and "the future," essentially lying to his customers.

He sells features that aren’t yet functional and lets his first customers help debug them. I consider this an unethical business practice—lying to customers is unacceptable.

For example, there’s a feature called "referential integrity." Suppose you terminate an employee who is the last person in a department—should you delete the department entirely, or keep an "empty department"? This is the kind of logic involved.

Ingres implemented this feature. At the time, Oracle’s approach was to include two pages of documentation in the manual explaining what referential integrity was (everyone agreed on this definition), but at the bottom of the page, it stated, “Not yet implemented.”

Host: I’ve spoken with people from Sun Microsystems, and their views on Ellison were similar. There’s also a saying that after Oracle acquired MySQL, many shifted to PostgreSQL, which helped make PostgreSQL a mainstream open-source database. So, what was the biggest change from Ingres to PostgreSQL?

Mike Stonebraker: The most fundamental change originated from an initial requirement. Back then, we wanted to support a GIS (Geographic Information System), which needed to handle data types such as points, lines, and polygons. But Ingres only supported standard types like integers, floating-point numbers, and strings, and could not efficiently support GIS—so we completely failed in that direction.

This has been on my mind.

Later, there was another example. Around 1985, relational databases introduced a datetime standard, and Ingres implemented the Gregorian calendar according to that standard. As a result, a customer called to say that you had implemented it incorrectly.

I said, "How could that be? We implemented it entirely according to the Gregorian calendar, and the date calculations are completely accurate." He replied, "That’s not what I need." He works in bond trading, where in his world, monthly interest is fixed regardless of whether a month has 28 or 31 days. In other words, his "date subtraction" rules differ from reality—for example, subtracting February 15 from March 15, he considers it 30 days. But in Ingres, these logic rules were hardcoded. He could only extract the data, perform the calculations at the application layer, and write it back, reducing efficiency by 2 to 3 times.

He asked me why subtraction couldn't be customized—that's exactly the issue. This is a scenario where you need "bonded time," just as you need points, lines, and polygons. That's why PostgreSQL was designed with an extensible type system: you can define arbitrary data types, and they run with high efficiency. This is the core idea behind PostgreSQL.

Of course, most business scenarios are well-served by standard data types, but as databases expand into more areas—such as abstract data types and stored procedures—extension capabilities become essential.

In addition, PostgreSQL supported inheritance (which AI researchers needed at the time) and "time travel" (historical data queries), though both were poorly implemented and later removed. Overall, however, it included a wide range of very interesting features.

Host: You mentioned you're very good at hiring top engineers. How do you identify these "exceptionally talented people"?

Mike Stonebraker: It’s usually obvious at a glance. I have a sense of “difficulty.” If a student completes three times the amount of work I consider reasonable, they are exceptional.

Host: You also said something interesting: “I can’t stand people who aren’t smart enough—it’s hard to communicate with them.” How do you determine if someone isn’t smart enough?

Mike Stonebraker: It’s simple—just talk to him for a while. Ask him technical details, like what his master’s thesis did, how it was specifically implemented, how error handling was managed, how many processes were used, and why threads weren’t employed—you’ll quickly be able to tell.

Host: You previously proposed the idea of "One size fits none," meaning that a one-size-fits-all database is not the optimal solution—it actually suits no one.

Mike Stonebraker: Yes, general-purpose database systems are not optimal. The idea of "one size fits all" often ends up fitting no one well. What you truly need is a database solution tailored to your specific requirements.

Host: So, which database products you see today still fall into this "one-size-fits-all" category?

Mike Stonebraker: When I wrote that paper in 2004, we happened to have an academic project at the time that later became StreamBase. Stream processing engines and relational databases looked completely different from each other. Meanwhile, we also had a general idea for columnar storage in data warehouses, which was later popularized by Vertica; columnar storage and row-based storage also appeared to be entirely different types of systems.

At that time, three vastly different implementations were already evident—each with little in common with the others—but all delivered performance an order of magnitude higher than traditional solutions in their respective scenarios. This alone speaks volumes: if your database system isn’t designed for your specific use case, you’ll immediately lose an order of magnitude in performance.

I still think this holds true today. For example, ClickHouse is a columnar store. Pinecone is also faster than solutions that force-fit user-defined types for text-based vector processing. So this remains valid today. I don’t see much difficulty in adding a unified parser on top of multiple different implementations. Yet PostgreSQL still hasn’t done this—it doesn’t truly implement columnar storage, so it lacks competitiveness in large-scale data warehousing scenarios. It also lacks multi-node support, which is now a basic requirement for large-scale data warehouses. So I believe this remains as true today as it was back then.

However, another equally true fact is that if you simply want to get started quickly and have a database requirement, the answer is usually still PostgreSQL. It has a large developer community, supports a wide variety of data types, is free, and it’s easy to find developers who know PostgreSQL, allowing you to get up and running quickly.

So I think it’s an excellent option for meeting basic, general-purpose requirements—as long as you’re not aiming for a million transactions per second or supporting a petabyte-scale data warehouse. In low-end scenarios, that “general-purpose solution” is definitely PostgreSQL; but in high-end scenarios, this conclusion no longer holds.

2 Once the index appears, the GPU struggles to be effective.

Host: Could GPUs bring new opportunities for database optimization?

Mike Stonebraker: Perhaps. But I think the biggest challenge is that GPUs are inherently SIMD—single instruction, multiple data—which directly conflicts with indexing.

If the index is the correct answer, the GPU is likely not a good answer.

In addition, you must design the entire system architecture to ensure that bandwidth from storage to computation does not become a bottleneck. If the GPU is merely attached as an add-on next to the CPU, the bus between the CPU and GPU often becomes the bottleneck itself.

Host: Can you explain why indexing performance degrades under SIMD mode?

Mike Stonebraker: For example, if I want to look up Ryan’s salary and I have a B-tree, I start by accessing the root node, find the split key that directs Ryan to one side, then follow the pointer down—that’s one definite memory access. Then I do it again, and again, typically repeating this three or four times.

This process is difficult to parallelize. So the answer is that indexing itself is not suitable for parallelization.

Host: You just mentioned B-trees. When you first implemented Ingres, were all of these components handwritten? I assume there weren’t any ready-made B-tree libraries available back then?

Mike Stonebraker: Yes, the earliest versions of Ingres were all written from scratch.

Host: What was the most difficult part to implement?

Mike Stonebraker: Query Optimizer.

Host: Why is it so difficult?

Mike Stonebraker: Because it’s genuinely hard. This component is algorithmically complex. Even today, if you ask any seasoned database programmer what the most difficult part of the system is, they’ll most likely say the optimizer.

Google chose the wrong direction; Amazon chose too many directions.

Host: After MapReduce emerged in the early 2000s, it rapidly swept through the entire data field. Many people were stunned, believing Google truly understood the field and that this was the most advanced technology available. But looking at your papers and views from back then, it seems you strongly disagreed. Why did you oppose MapReduce so vehemently?

Mike Stonebraker: At the time, many people who didn’t really understand things assumed that Google was smart and must know what they were doing, so we just followed their lead. As a result, everyone started working on Hadoop or trying to align with Hadoop’s approach.

But Hadoop is ridiculously inefficient.

At the time, people like Dave DeWitt and others who contributed to our 2011 paper all understood distributed databases and knew that a distributed database system could easily outperform Hadoop. Our 2011 paper essentially made this point—and history has since proven us right.

But Google has done more foolish things than just this one.

At the time, they also believed that eventual consistency was the correct approach to concurrency control—a perspective that Google actively promoted during that period. But this was fundamentally wrong. Everyone in the database field was saying, “You’re insane,” because it only addresses a very specific type of problem, one that is actually rare in the real world.

Host: Why would they pursue eventual consistency?

Mike Stonebraker: Imagine you have a database on the East Coast and a database on the West Coast, each serving as a replica of the other, and you want to keep them in sync.

If I now want to perform a transaction that reduces the inventory of a certain product in the West Coast warehouse by one, I must first synchronize this update to the East Coast before confirming that the update was successfully applied there. Then, to ensure the entire transaction is truly completed, I need another round-trip communication to verify that both sides have committed correctly. Distributed commits are this expensive—and they still are today.

So someone thought, “What if I just reduce the inventory by one on the West Coast first, then asynchronously send a message to the East Coast without including it in the transaction—so the East Coast will eventually reduce by one as well?”

Conversely, if an item is also sold on the East Coast, it asynchronously sends a message over, and eventually both sides gradually converge to consistency.

The problem is that if the system allows inventory to go below zero, it could result in a scenario where someone on the East Coast and someone on the West Coast simultaneously sell the last item. As a result, the system’s inventory would become negative one, meaning one person ultimately won’t receive the product.

If you were like Amazon and could say “usually shipped within 24 hours,” you might be able to tolerate some degree of overselling. But most businesses can’t. That’s why eventual consistency simply doesn’t work.

We mentioned referential integrity a long time ago. Similarly, sales systems have analogous integrity constraints, such as “inventory must be greater than or equal to zero.” In such cases, eventual consistency will fail.

Later, Jeff Dean of Google finally recognized this issue. When developing Spanner, they used a traditional transaction system. In other words, Google completely abandoned eventual consistency and also abandoned MapReduce.

Host: So, is the fundamental trade-off essentially exchanging accuracy for performance?

Mike Stonebraker: Yes, it’s a trade-off between performance and data integrity. If you don’t care about your data at all, then you can certainly accept these poor consequences.

Host: When Google was doing these things you clearly thought were wrong, did you communicate with their team?

Mike Stonebraker: We reached out to them before our 2011 paper. We asked if they’d be interested in collaborating on something, but they had no interest and declined outright.

Host: Besides Google, have you seen similar database solutions at other major tech companies that you clearly disagree with, such as Amazon or Facebook?

Mike Stonebraker: I gave a talk at Amazon about three years ago, during which I outlined everything I thought they were doing wrong.

I think the problem with Amazon is that they support around 15 different database systems, which is probably about 12 too many.

I think it has to do with their own culture. I told them at the time that they were supporting too many database types, but to this day, they haven’t decided to eliminate any of them.

Host: Why do you think 15 should be reduced to 3?

Mike Stonebraker: Because they support graph databases, and it’s long been known that graph databases are almost never the most performant solution.

If you prefer a user interface with nodes and edges, that’s perfectly fine. You can easily add a layer on top of a relational database to provide this user model.

Many current database systems are duplicating efforts that other databases already perform better. Therefore, the conclusion is that databases with insufficient performance and markets too small to support their maintenance costs should be phased out.

Host: You’ve had a tremendous impact on the entire industry from an academic perspective. I’ve always been curious: why haven’t you gone straight into industry? For example, why not join a company like AWS as a senior distinguished engineer—you could still make a huge impact that way. Why do you prefer to stay in academia and contribute in this way?

Mike Stonebraker: Because then you’d have a boss. You’d be bound by company rules, restricted in publishing papers, limited in speaking at conferences, and even restricted from researching what competitors are doing, since companies often won’t let you discuss many things publicly—especially not anything competitors don’t want revealed.

But more importantly, I really enjoy being at a startup. After the commercial version of Postgres was acquired by Informix, I worked part-time at Informix. It was a 2,000-person company where I felt completely unable to make an impact, as it was filled with bureaucracy and essentially ran by whatever the CEO wanted.

So I think I’m probably not suited for that kind of environment. I’m not good at office politics, and I struggle to interact with people I don’t find intelligent. Ultimately, I just don’t mesh well with large corporations.

4. Replace the upper part of Linux with a database.

Host: I’d like to talk about Deboss. I think it’s an interesting technical model. Could you explain what exactly Deboss is?

Mike Stonebraker: This academic project began around 2019 or 2020. Its origins are closely tied to Matei Zaharia, who was then a professor at Stanford, a co-founder of Databricks, and the original creator of Spark.

He said that at the time, Databricks was essentially running users' Spark jobs on the cloud, and at any given moment, they might be scheduling millions of Spark tasks simultaneously. Therefore, they had to build a scheduler capable of deciding "who runs next" at a scale of millions.

He said they tried various schedulers written by people in the operating system field, but none scaled well. In the end, they put all scheduling data into a PostgreSQL database, effectively using a PostgreSQL application to handle scheduling.

This immediately made us realize that, at its core, most tasks in an operating system are essentially about managing large volumes of data—and such tasks should naturally be handled using database technology. So our thinking shifted to: Why not replace at least the upper half of Linux with a database system?

This was the core of that academic project. In the early 2020s, we worked on it with Berkeley and Stanford, and the results were highly successful, clearly demonstrating that this path is viable.

During this process, Stanford also extended JavaScript, since you need a programming environment that can communicate with the underlying implementation. If you're running a programming language on top, and the underlying "operating system" is essentially a database, the most natural approach is to store all state within the database—and that's exactly what they did.

So we ended up with both a new programming language model and a new operating system model. The natural next question was: could we turn this into a company? Later, when we approached venture capitalists, almost everyone gave us the same response: if you think you can replace Linux, you’re dreaming. But your programming language stuff is quite interesting.

Since we’re essentially building a JavaScript extension that enables any program to naturally inherit many advantages of database systems—such as persistent state, transaction support, and automatic failover after failures—we secured funding in 2023 and founded Deboss. The project had always been called Deboss, so we kept the same name for the company. However, what we’re doing has essentially evolved into something closer to a programming language business.

Deboss now offers a set of TypeScript, a set of Java, a set of Go, and a set of Python, all of which are nearly seamless. What you write looks just like ordinary native code.

In the cloud era, nearly every factor pushes you to organize your applications as workflows. So we decided to directly support workflow systems. In Deboss’s workflow models across these four languages, each step—each small micro app, whatever you call it—is transactional. The entire workflow is persistent, meaning once a step is completed, it is not forgotten.

Moreover, it’s clear that if there is market demand, we can also make the entire workflow atomic—meaning the process either completes fully or appears as if it never happened. This gives it many excellent properties, and it’s significantly faster and much easier to use than existing competitors. The company is currently continuing to sell products and innovate in this direction.

In short, you need to make the application’s state persistent by storing it in a database, and then figure out how to make it fast. I think their current business model aligns closely with what we discussed earlier—first attracting the developers who directly use these tools at the most fundamental level.

What we’ve been doing is asking frontline developers what they’re missing and what we don’t yet offer, then quickly filling those gaps and convincing them to try it. We’ve been very successful with startups, since these companies only want the best tools. Now we’re also gradually making inroads into larger enterprises.

This is an interesting market. So far, about two-thirds of clients are working on agentic AI—meaning they have a large language model surrounded by various components that provide it with additional signals.

However, most current agentic AI systems are still read-only. For example, if you want to predict whether Ryan will be a good customer, the system simply runs a series of logic checks and outputs a result for someone else to act on. Fundamentally, it is read-only and does not actually modify Ryan’s credit score or similar data.

But I believe that soon, the entire world will shift toward using agents to handle read-write applications. Once we reach that point, these systems will become highly “database-like,” and Deboss is exceptionally skilled at handling exactly this kind of task.

For example, suppose you want to write one or two agents to transfer $100 from my account to yours. Those agents must first deduct $100 from my balance and then add $100 to your balance, and both agents must agree on the “commit” — otherwise, the entire transaction must be rolled back.

In other words, this workflow must be atomic, as I mentioned earlier—either all of it happens, or it appears as if it never happened.

So I believe demand in this market will continue to evolve, with more people increasingly wanting these systems to not only read but also write. I think this is beneficial for the entire market and also for Deboss.

Host: So the product now available to developers in the market is quite different from the original research project. The original project truly aimed to replace the core internal state of the operating system with a database. I have to say, that’s really cool—I never thought it was possible to put all of an operating system’s state into a database. But there must be some trade-offs involved, right?

Mike Stonebraker: Actually, there’s no significant cost. A file system built on top of a DBMS is faster than the Linux file system. The performance of the scheduling engine is on par with other scheduling engines. You can also make everything fault-tolerant, meaning you essentially get high availability without having to do much extra.

So the answer is, there are virtually no drawbacks.

Host: Then why doesn’t Linux just absorb this system and upgrade it internally?

Mike Stonebraker: They should theoretically do this. In other words, the low-level tasks like device drivers and similar messy work should obviously be retained, since there’s too much of it and no one wants to rewrite it. But everything else should be replaced with database-style implementations.

Host: Have you mentioned this to people in the Linux community? What’s their usual reaction?

Mike Stonebraker: During the academic project phase, whenever I shared this idea with people in the operating systems field, they felt threatened, believing that the database crowd was coming to take over their territory.

People in the programming language field feel much the same—how, are you now saying that the runtime of a programming environment should also be implemented using a database?

Host: That's interesting. If it's technically correct, it may eventually take over existing solutions.

Mike Stonebraker: Java took 10 years to gain widespread adoption. I just think the time constant for这种事情 is inherently large.

5. In a real data warehouse, the accuracy rate of LLM-generated SQL is 0%.

Host: We’ve talked a lot about the history of databases. I’m also curious—what do you think are the unresolved challenges in the database field, and what will the future look like?

Mike Stonebraker: Here, I’d like to discuss two things. First, like others, about three years ago, we began exploring what large language models are truly suited for.

We have been trying to apply what is currently called text-to-SQL to real-world databases, particularly real-world data warehouses. We have tested this technology on four live production data warehouses. We obtained the actual workloads from these systems—the queries that real users have run—and then reverse-engineered the natural language descriptions corresponding to those SQL queries.

So we now have four sets of benchmarks, each containing both text and SQL.

Host: When you mention text-to-SQL, do you mean using natural language to prompt the model? For example, by simply saying a sentence in English?

Mike Stonebraker: Yes, like "find everyone over the age of four" or "tell me all the MIT professors who have won the Turing Award." By that logic, large language models should be very good at this.

Among the publicly available text-to-SQL benchmarks, there is one called Spider and another called Bird. The best large language model systems perform quite well on these benchmarks, achieving accuracy rates of around 80% or higher. While not superhuman, this is already very impressive and reaches a level where you’d seriously consider deploying it in practice. Current leaderboard results are around 85% accuracy, meaning you might feel it’s not quite ready for prime time yet, but it’s very close.

However, on our benchmark, the accuracy of large language models is 0%. If you add various enhancement techniques like RAG, the accuracy can reach about 10%. If you directly provide the FROM clause in the prompt—explicitly telling it which tables to access and which join conditions to use—the accuracy can reach approximately 35%.

Therefore, from the standpoint of whether it can be deployed in production, this technology is still far from adequate, and there is no foreseeable prospect in the near term that it will ever reach that stage—perhaps it never will.

Host: What exactly is the difference?

Mike Stonebraker: First, the data in data warehouses is not included in the training corpora of large models. Large language models are primarily trained on publicly available data, while the real business data in data warehouses is not present at all. There’s an old saying that holds true: if a model has never seen this data before—or at least hasn’t seen it multiple times—it’s highly unlikely to generate it correctly. That’s the first point.

Second, queries in benchmarks like Spider and Bird typically consist of about 10 to 20 lines of SQL. However, in real-world data warehouses, SQL queries often start at 100 lines or more, making the complexity orders of magnitude greater.

Third, the schemas for Spider and Bird are very clean: table names are intuitive, column names are clear, and there is no redundancy. Data warehouses, however, are not like this. They often contain numerous materialized views, resulting in significant redundancy; column names are frequently filled with underscores, abbreviations, and obscure naming conventions that make their meanings nearly impossible to guess.

This makes the problem much more difficult. Additionally, real-world systems contain vast amounts of highly “localized” and highly specialized data. For example, MIT has a “J-term,” referring to a one-month term in January. This is not unique to MIT, but it is also not common or widespread enough to appear in training corpora.

So, the data isn’t in the training set, the queries are more complex, the schema is a mess, and there’s a pile of system-specific data on top—all of which together make this whole thing completely unworkable. And every data warehouse I’m aware of basically shares these characteristics. So my assessment is that this technology simply doesn’t work now, and it won’t work in the near future.

Host: So what are you doing now?

Mike Stonebraker: First, we released our own benchmark, called Beaver. It is a de-identified and abstracted version derived from four real data warehouses. So, if anyone truly believes they excel at text-to-SQL, let them run a legitimate benchmark instead of relying on fake ones to feel good about themselves.

Second, given the issues I just mentioned, if you don’t have a JOIN condition or a FROM clause, you’re already out of luck. Furthermore, if you can’t break down a complex query into simpler parts, you’re still out of luck.

So this makes me think that the most reasonable approach is to first transform the input retrieved by the system into simpler segments, each clearly containing the FROM clause and JOIN conditions. This is the first point.

Second, once you need to work with two different structured databases—such as a data warehouse and a CRM system—using a large model to perform joins between structured data is, in my view, a bad idea. It’s far more reliable to keep them as tables and use SQL to perform the joins.

So our current approach is to turn everything into tables. For example, we’re currently collaborating with the Munich city transportation department in Germany, which has six full-time staff members dedicated to responding to citizens’ complaints and inquiries. The questions vary widely—such as, “Why is the green light at the intersection near my house not long enough for me to cross?” “Why doesn’t the tram stop long enough for me to board?” “Why does this tram only come once an hour?”—all kinds of these types of questions.

And the data sources behind them are highly diverse: train schedules are in SQL, traffic light timing is in SQL, intersection information is in CAD, German federal traffic regulations are in text, and Munich’s own regulations are also in text. In other words, you need to correlate SQL, SQL, CAD, text, and text—these five types of data.

Our approach is to convert everything into SQL, or essentially into tables, and then perform joins using a method that closely resembles a query optimizer. That’s exactly what we’re doing now. I believe others may have different approaches, but I think this direction holds tremendous potential because everyone is truly eager to make this happen. That’s the first thing.

The second thing is what we discussed earlier—agentic AI. Once it moves from read-only to read-write, it immediately becomes a distributed database problem. You need atomicity, consistency—all those old issues come back. I think this is also a very interesting direction. So right now, that’s primarily what I’m working on.

Host: On your benchmark, large models are currently at 0%. So, what level can humans typically achieve? For example, what average performance can a person who truly understands SQL reach?

Mike Stonebraker: As long as you eliminate ambiguities in natural language first, a programmer familiar with SQL and able to read schemas will achieve very high accuracy.

Host: For example, more than 90%?

Mike Stonebraker: Yes, at least that scale.

Host: Understood. To be honest, I’m still surprised that large models perform so poorly on this benchmark. Maybe after this episode airs, someone from Anthropic will reach out and say, “Let’s give it a try.”

Mike Stonebraker: I’m curious to see how far they can go, because it’s actually a great opportunity for those who truly want to understand databases deeply.

6 Computer science may no longer be a growth industry

Host: If someone wants to systematically learn about databases, what textbooks or research papers would you recommend?

Mike Stonebraker: Joe Hellerstein and I co-authored a book called "Readings in Database Systems," often referred to as the "red book." Although it’s now eight years old, I still think it’s an excellent entry point for classic papers from eight years ago and earlier. For more recent developments, look into the influential and widely cited papers that have emerged in the database field since then.

Host: If you could go back to the time right after graduation and give your younger self some advice with today’s knowledge, what would you say?

Mike Stonebraker: When I first arrived at Berkeley to teach, I barely thought about it and said, “Let’s build a database system.” But at the time, we knew almost nothing about databases, nothing about implementation, and our programming skills weren’t anywhere near as strong as Bill Joy’s. So, jumping straight into such a crazy endeavor was, in itself, extremely crazy.

That’s just how people are—start by diving in headfirst, learn as you go, and build things while learning. So I believe the answer is to think outside the box, dare to imagine wild ideas, and then actually go do them.

But for me, a more relevant question right now is: If you were just starting out today, what would you choose to study?

Because I feel that computer science may no longer be a growing field in the future. I’m not sure I’d still recommend that an 18-year-old today study computer science.

Healthcare and various trades such as construction and repair still appear to be relatively safe options; many other fields carry significantly higher risks.

Of course, if you're close to earning your PhD and are considering what to do next, things become much simpler: go for the most prestigious position available, find a mentor who’s willing to guide you, and choose a direction that doesn’t follow the crowd. For example, what we’re doing with Rubicon is definitely not following the mainstream. So, find an unconventional path and figure out how to make it work.

My wife and I used to tell our children: “Follow your passion, and the money will take care of itself.” But to be honest, I didn’t believe it at all. Still, I think you can only tell your children—and your grandchildren—that same thing.

Host: If you don’t believe this statement yourself, why are you saying it?

Mike Stonebraker: My wife is a great example. She earned both her bachelor’s and master’s degrees in computer science, but what she truly wanted to do was become a teacher—specifically, an elementary or secondary school teacher. However, her parents told her at the time that it wasn’t feasible because the income was too low.

I think she has regretted that decision ever since. She wasn’t truly passionate about computer science—it was just a means to make a living.

So I think you should still pursue something you're truly passionate about. At least you won’t go hungry. You might not make a lot of money, but you’ll likely be much happier than doing something you have no passion for.

Because I know many people who see their jobs as just jobs, with real life only happening between 5 p.m. and 8 a.m. the next day. But I don’t feel that way at all—I genuinely love what I do. Whether I make a lot of money or not, that won’t change.

Original video link: https://www.youtube.com/watch?v=YPObBOwIrHk

This article is from the WeChat public account "InfoQ" (ID: infoqchina), authored by Tina.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.