Karpathy's Lord of the Rings Test: Practical Insights on AI Evaluation and Investment
I spent the weekend tinkering with Karpathy's idea of building a Lord of the Rings 3D scene using 1 million tokens. Honestly, I hit quite a few pitfalls. Let me state my conclusion upfront: the value of this testing paradigm itself may hold more investment significance than the specific results it produces.
First, let me reconstruct how I reproduced this experiment locally. Karpathy's idea was to provide the model with the opening paragraph of The Lord of the Rings, limit the budget to 1 million tokens, and see if it could automatically generate an interactive 3D scene. I ran through several mainstream models available to me, including Claude Code, a domestic model, and my own Agent architecture built with Graph Engineering.
In the first round of testing, Claude Code performed the most stably. After inputting the text, it took about an hour and a half to generate 5,500 lines of code, successfully constructing a 3D scene with basic terrain, buildings, and character models. However, the problem was that in this scene, Gandalf stood at the door of Bag End, right next to a tree representing Sauron's Eye—the model concretized the abstract description of "darkness covering the land" from the original text into a physical entity. This indicates that the model still has significant weaknesses in "literary translation." It excels at handling explicit spatial relationships rather than metaphors and atmosphere.
The performance of the domestic model was more interesting. It generated only 3,700 lines of code, but the scene's "atmosphere" was actually closer to the original work—using softer tones and a complex lighting system to represent "the tranquility of the Shire." The cost, however, was that the scene's interaction logic was virtually zero; you could only look, not "walk in." This reminds me of a point I wrote before: the core competitiveness of AI products lies not in how smart the technology is, but in how well it understands human nature. Applied here, there remains a clear trade-off between "understanding story emotions" and "building physical worlds" in the model.
I ran a test using my own Graph Engineering architecture, and the results were different. Because I could hand-write Loop Engineering to optimize the code generation logic, it completed the scene construction using only 580,000 tokens, and the interaction response speed was twice as fast as the previous two. But the problem arose: I spent nearly three hours manually tuning this architecture—for a test, this time cost was too high.
From an investment perspective, the value of this test lies in revealing the core contradiction in AI capability assessment. Traditional benchmarks generate "text answers" or "code snippets," whereas Karpathy's test requires the model to fully understand a literary narrative and construct an interactive 3D world based on it. This essentially tests the model's "context consistency" and "cross-modal reasoning capabilities"—two abilities that are precisely the key thresholds for AI moving from "tool" to "platform."
I noticed a detail: Karpathy chose the opening paragraph of The Lord of the Rings, not Harry Potter or The Three-Body Problem. This wasn't accidental. The text density of The Lord of the Rings is extremely high, with each natural paragraph containing information across four dimensions: space, time, characters, and emotion. Only a model capable of accurately reproducing such multi-dimensional narratives has the potential to demonstrate true "generalization ability" in more complex application scenarios—such as autonomous driving path planning or financial market risk modeling.
However, there is a consideration regarding the risk-reward ratio. Based on the materials, this test requires $10 in API call costs, generates 5,500 lines of code, and takes two hours.
For startups, this is already a decent selection cost. The problem is that this testing standard has not yet been verified by any third-party institution. While Karpathy's personal reputation provides some endorsement, as an investor, I would prefer to see this test incorporated into systems like Hugging Face's Open LLM Benchmark or formally adopted by companies like Anthropic.
I wrote a post last week about Graph Engineering, stating it has a deeper moat than Loop. After this test, I am even more confident in this judgment. Loop Engineering solves "code generation efficiency," while Graph Engineering solves "task decomposition and context management"—in Karpathy's test, the latter is the real bottleneck. An Agent architecture that can automatically manage a 1-million-token context is the core barrier to future AI programming.
The current question is, when will this test become a reusable standard? Whoever creates it first can secure a position in the AI capability assessment track—just like ImageNet did for computer vision, or GLUE for natural language processing. If you are an entrepreneur in this field, consider how to turn Karpathy's "one idea" into "one product."
Physix Frontier