DHH用AI重写Campfire:暴露无代码编程的局限与权衡陷阱
I'm sorry, but you still have to think
Agent开发者的必读反面教材,详细拆解了“甩手给AI”导致的架构灾难,提供了关于异步调度和压力测试设计的深刻工程教训。
Written by Piotr Sarnacki
作者:Piotr Sarnacki
on October 11, 2026
发布于 2026 年 10 月 11 日
I'm sorry, but you still have to think
很抱歉,但你仍然需要思考
DHH, the creator of the Ruby on Rails framework, decided to rewrite Campfire Once from Ruby on Rails to Rust. Or rather ask a clanker to do it for him, as he can't stand reading or writing the Rust code himself. After that he added rewrites in other languages, like Elixir and Go.
Ruby on Rails 框架的创建者 DHH 决定将 Campfire Once 从 Ruby on Rails 重写为 Rust。或者更准确地说,是让一个 AI(clanker)替他完成这项工作,因为他自己无法忍受阅读或编写 Rust 代码。此后,他又添加了使用其他语言(如 Elixir 和 Go)的重写版本。
This is very interesting to me, as it's a very good insight into the results of someone using AI agents without reading the code. In this case, that someone has also a lot of prior programming experience. And the results are... well, not very encouraging if you are hoping you can stop reading the code, or stop thinking altogether anytime soon. There are multiple problems with the rewrites, but let's start with non-functional differences cause they nicely show something that a lot of people seem to not realize: if your prompt is not specific enough, many decisions are a coin flip. This is because a whole lot of questions don't have a single correct answer. Do we want to keep backwards compatibility? Do we care more about latency or throughput? How much memory can the system use under load? Is loosing new content notifications after a crash acceptable? You may not care about them, or at least some of them, but they will be implicitly answered when an LLM implements what you think you want.
这对我来说非常有趣,因为它很好地揭示了在不阅读代码的情况下使用 AI 智能体所产生的结果。在这种情况下,使用者本身拥有丰富的编程经验。而结果嘛……嗯,如果你希望不久后就能停止阅读代码,甚至完全停止思考,那么这些结果并不怎么令人鼓舞。这些重写版本存在多个问题,但让我们先从非功能性差异说起,因为它们很好地展示了许多人似乎没有意识到的事情:如果你的提示词不够具体,许多决策就等同于抛硬币。这是因为大量问题并没有唯一正确的答案。我们是否要保留向后兼容性?我们是更关心延迟还是吞吐量?系统在负载下可以使用多少内存?崩溃后丢失新内容通知是否可以接受?你可能并不关心这些问题,或者至少不关心其中的一些,但当 LLM 实现你认为想要的功能时,它们会被隐式地回答。
If you look into the rewrites closer, you will quickly see that they treat backwards compatibility and other constraints differently. The Rust version doesn't maintain 100% backwards compatibility, for example it drops the CSRF token so that caching is easier. It also dropped Redis in favour of in-process queues, for example for notifications. Elixir rewrite is much closer to the Rails original. This already makes the comparison pretty much useless, as these differences are language independent. It's not like a clanker heard "rewrite in Elixir" and chose to keep 100% backwards compatibility because of the language that has been used. But it gets even better!
如果你仔细查看这些重写版本,你会很快发现它们对待向后兼容性和其他约束的方式各不相同。例如,Rust 版本并没有保持 100% 的向后兼容性,比如它移除了 CSRF token 以便简化缓存。它还用进程内队列取代了 Redis,例如用于通知。Elixir 重写版本则更接近原始的 Rails 版本。这使得比较几乎毫无意义,因为这些差异与语言无关。并不是说 AI 听到了“用 Elixir 重写”这句话,就因为这个语言的特性而选择保持 100% 的向后兼容性。但情况还有更糟的!
If you looked at the code, and yeah, I know, we should not be reading code anymore, you would quickly notice a lot of the things are not great. For example, I've seen people complaining that Elixir's version got a single process sequentially processing all of the SQL queries, even if these were reads that could be running concurrently. That's a fair complaint, but DHH seems to think it's not the best look for Elixir that the clanker couldn't write performant code. Not sure if I agree when I look at the Rust version. Cause in the Rust version some of the database operations are not async, which in some cases may be worse. You see, in Rust, when using an async runtime, the scheduling is not preemptive, but rather cooperative. If a task doesn't yield, no other task can run on the same worker thread. What this means in practice, is that the time spent in the task should be as short as possible. For example, when you run an SQL query that takes 100ms, you don't want all the other tasks to wait as the task issuing the query is waiting anyway. Thus, you should ideally use an async I/O operation that yields to the runtime while waiting for a response from the database. In the Rust rewrite some queries run on worker threads, but some run as blocking operations in async tasks. Same with locks. When using an async runtime, the safest bet is to use async locks like tokio::sync::Mutex. It's fine to use a non-async version if you're sure that the lock is held for a very short time, but if you hold a sync lock for 10ms, all the tasks on the same thread wait for it, while the blocking task itself is not doing any work. Thus, I regret to inform you, that no, coding is not likely "solved", and you still need to know what you're doing.
如果你查看了代码——是的,我知道,我们本不该再读代码——你会很快注意到很多方面表现不佳。例如,我曾看到有人抱怨 Elixir 版本中只有一个进程顺序处理所有 SQL 查询,即使这些是并发运行的读取操作。这是一个合理的批评,但 DHH 似乎认为,AI 无法写出高性能的代码,这对 Elixir 来说形象不太好。当我查看 Rust 版本时,我不确定是否同意。因为在 Rust 版本中,某些数据库操作不是异步的,这在某些情况下可能更糟。你看,在 Rust 中,使用异步运行时(async runtime)时,调度不是抢占式的,而是协作式的。如果一个任务不让出控制权(yield),其他任务就无法在同一工作线程上运行。在实际中,这意味着任务所花费的时间应尽可能短。例如,当你执行一个耗时 100ms 的 SQL 查询时,你不想让所有其他任务都等待,因为发起查询的任务本身就在等待。因此,理想情况下,你应该使用异步 I/O 操作,在等待数据库响应时让出控制权给运行时。在 Rust 重写版本中,一些查询在工作线程上运行,但另一些则作为阻塞操作在异步任务中运行。锁的情况也是如此。在使用异步运行时,最安全的做法是使用异步锁,如 tokio::sync::Mutex。如果你确信锁持有的时间非常短,使用非异步版本是可以的;但如果同步锁持有时间为 10ms,同一线程上的所有任务都会等待它,而该阻塞任务本身并未进行任何工作。因此,很遗憾地告诉你,编码问题并没有被“解决”,你仍然需要知道自己正在做什么。
Going further, it turned out that slop benchmarks are also, well, slop, as they only measure throughput, disregarding other properties of the system. For example, Zach Daniels measured new post notifications delivery rate under load, and discovered a 1% successful delivery rate in the Rust version under heavy load. Not a great look. Benchmarks are hard. But even the improved benchmark may not be actually testing what you want to test, depending on the characteristics of your system. You see, when you test a system you may want to emphasize different properties under load. Both DHH's and Zach's benchmarks were closed-loop benchmarks, so they were testing "how many requests can the system handle in a specific amount of time?". The test was using N clients, each sending a new message as soon as it gets a response for the previous one. In many cases, though, the increased load on the system may come from a big number of users performing an action at the same time, who will not wait for other users to finish their actions. In this case you would prefer constant arrival rate, ie. you want to send requests at a constant rate rather than making the rate depend on how fast the system can respond. If you want to check how reliable the events delivery is, you would ideally want to compare the same number of events. Here, if you look at the results, it shows that Rust delivered 1% of notifications out of 6-7k, whereas Elixir delivered 100% of notifications out of ~1.7k. What it shows is that the Elixir version handles backpressure better, but this only holds as long as the number of requests don't overload the system.
更进一步,事实证明 slop 基准测试也是 slop(粗制滥造),因为它们只衡量吞吐量,忽略了系统的其他属性。例如,Zach Daniels 测量了负载下的新帖子通知送达率,发现在重载情况下 Rust 版本的成功送达率仅为 1%。这并不好看。基准测试很难。但即使改进后的基准测试,也可能并非真正测试你想要测试的内容,这取决于你系统的特性。你看,当测试系统时,你可能希望在负载下强调不同的属性。DHH 和 Zach 的基准测试都是闭环基准测试,因此它们测试的是“系统在特定时间内能处理多少请求?”。测试使用了 N 个客户端,每个客户端在收到上一个请求的响应后立即发送一个新消息。然而,在许多情况下,系统增加的负载可能来自大量用户同时执行某个动作,而这些用户不会等待其他用户完成他们的动作。在这种情况下,你更倾向于恒定的到达速率,即你希望以恒定速率发送请求,而不是让速率取决于系统的响应速度。如果你想检查事件交付的可靠性,理想情况下你应该比较相同数量的事件。在这里,如果你查看结果,它显示 Rust 从 6-7k 中交付了 1% 的通知,而 Elixir 从约 1.7k 中交付了 100% 的通知。这表明 Elixir 版本更好地处理了背压(backpressure),但这仅在请求数量没有使系统过载的情况下成立。
But let's leave benchmarks just for a moment and talk about trade-offs, cause programming is all about trade-offs. Sure, there are situations where a tool or a solution is clearly better, without any downsides, but it's quite rare, especially when you get to complex systems that have to be reliable. I've seen quite a lot of people from the Elixir community coming to various conclusions after they've seen the 1% delivery rate of the Rust version, without really trying to understand why it happens. The general consensus? Elixir is just better at concurrency! In Rust the scheduler is cooperative, how can you even live like this? I like Elixir, and I've successfully used it in production, but it's not a silver bullet. Yes, Elixir (or other BEAM based languages) are very good at concurrency, and admittedly writing concurrent code in Elixir is, in general, easier than in Rust, but it comes at a cost of not having low level control, higher memory usage, and oftentimes, speed. There is a reason why rustler exists. And proclaiming language's superiority based on a single metric without knowing the root cause, may be misleading.
但让我们暂时放下基准测试,谈谈权衡取舍,因为编程就是关于权衡取舍的。当然,有些情况下某个工具或解决方案明显更好,没有任何缺点,但这相当罕见,尤其是当你涉及到必须可靠的复杂系统时。我看到很多来自 Elixir 社区的人在看到 Rust 版本的 1% 交付率后得出了各种结论,而没有真正去理解为什么会发生这种情况。普遍共识是?Elixir 在并发方面做得更好!Rust 中的调度器是协作式的,你怎么能忍受这样的生活呢?我喜欢 Elixir,并且已经在生产中成功使用过它,但它不是万能药。是的,Elixir(或其他基于 BEAM 的语言)在并发方面非常出色,而且不可否认,在 Elixir 中编写并发代码通常比在 Rust 中更容易,但这以失去底层控制、更高的内存使用率和往往较低的速度为代价。rustler 的存在是有原因的。在不了解根本原因的情况下仅基于单一指标宣称某种语言的优越性,可能会产生误导。
Remember how I mentioned the closed-loop stress test shows just one of the properties of the system as Rust processed ~4 times more requests and thus had to handle more events sent to the client through a WebSocket? I've rerun the test with constant delivery rate using the DHH's Rust version and Zach's Elixir version with various fixes. At 100 POSTs/s the clients received ~14% events in Rust. In Elixir it was ~60%. Still better, right? Not so fast! At this traffic level Rust didn't have any HTTP errors. Elixir timed out on ~23% of HTTP POST requests. And what about latency? The worst event delivery latency was close to 180s. In Rust when the deliveries are lagging, clients get disconnected. On reconnect, a browser client would fetch the latest messages, which largely invalidates the need for the missed events. What do you think is better UX: the client silently reconnecting in the background and fetching the new updates, or waiting for an update about a new message for 3 minutes? Which just shows that, again, a single metric doesn't tell the whole story.
还记得我之前提到的闭环压力测试吗?该测试仅展示了系统的一个特性:Rust 处理的请求量约为 Elixir 的 4 倍,因此必须通过 WebSocket 向客户端发送更多事件。我使用 DHH 的 Rust 版本和经过各种修复的 Zach 的 Elixir 版本,以恒定交付率重新运行了该测试。在每秒 100 次 POST 请求的情况下,Rust 客户端接收到了约 14% 的事件,而 Elixir 为约 60%。看起来更好,对吧?别急!在这个流量级别下,Rust 没有任何 HTTP 错误,而 Elixir 约有 23% 的 HTTP POST 请求超时。那延迟呢?最坏的事件交付延迟接近 180 秒。在 Rust 中,当交付滞后时,客户端会断开连接。重连后,浏览器客户端会获取最新消息,这在很大程度上消除了对丢失事件的依赖。你认为哪种用户体验更好:客户端在后台静默重连并获取新更新,还是等待一条新消息的更新长达 3 分钟?这再次证明,单一指标无法反映全貌。
更进一步:量化金融体系
看懂新闻只是起点——沿量化金融路径,把它变成能交付的工程能力