In a hands-on technical experiment reported by daily.dev, a computer science student tested three agentic AI coding tools—Claude Code, Codex, and Google Antigravity—to build an open-ended, feature-complete note-taking app from scratch. The primary finding: Codex, powered by OpenAI’s GPT-6 models, emerged as the clear winner by delivering a polished and cohesive application named Folio, outperforming rival implementations in architectural focus and user interface execution.
Building Folio Under Strict Tool Constraints
The experiment required each AI coding tool to build a modern, fast, and feature-rich note-taking app based on an extensive, open-ended prompt without any generative AI features, chatbots, or LLM integrations inside the final product. The shared feature brief demanded normal page notes, infinite canvas notes, handwriting and stylus support, typed text mixed with handwriting, fast note switching, a Home and Recent view, universal search, quick capture, temporary scratch notes, research notes handling PDFs and screenshots, bidirectional backlinks, PDF import and annotation, audio recording, version history, file exports, and standard notebook organization.
To keep the test fair, fresh sessions were initialized for all three platforms. Codex kicked off its build using GPT-6 Astra at high reasoning levels. As session constraints approached, the tool transitioned to GPT-6 Luna at medium reasoning. Codex was the only platform in the test to exhaust a five-hour usage limit, spending that compute time rigorously testing note persistence, search indexes, and canvas export workflows.

The resulting application, Folio, featured an aesthetically appealing and highly usable interface. It rigorously adhered to the prompt’s structural and practical requirements without getting distracted by unrequested additions.
Divergent Results Across Claude Code and Google Antigravity
While Codex secured the top spot, the other two contenders yielded sharply contrasting development trajectories and output metrics. Claude Code, running on Claude Opus 5.5, generated a massive volume of code—reaching approximately 489,000 output tokens alongside 75.3 million prompt-cache reads. In stark contrast, Codex consumed roughly 82,000 output tokens for the exact same software engineering task.
Despite Claude Opus 5.5 generating over six times the output volume, its resulting application—named Margin—felt less cohesive and polished than Codex’s Folio. The disparity stemmed largely from Claude Code over-engineering unrequested features and building out auxiliary systems outside the requested scope, complicating the overall user interface.
Google Antigravity presented a different set of engineering hurdles during the session. The platform declared the project finished prematurely, leaving out several explicitly requested features such as PDF annotation, audio recording, and backlink parsing, and categorizing them as future work. Daily.dev reported that Antigravity required manual developer intervention to resolve Tailwind CSS migration issues and TypeScript compilation errors before the application could execute properly. Once flagged by the user, however, Antigravity implemented corrective fixes.
| AI Coding Tool | Flagship Model Used | Output Token Volume | App Name | Development Outcome |
|---|---|---|---|---|
| Codex | GPT-6 Astra / Luna | ~82,000 tokens | Folio | Winner; feature-complete, highly polished UI, rigorous workflow testing. |
| Claude Code | Claude Opus 5.5 | ~489,000 tokens | Margin | Over-engineered unrequested features; less cohesive UI despite massive output. |
| Google Antigravity | Proprietary Agentic Setup | Not specified | Unnamed | Prematurely declared completion; required Tailwind and TypeScript fixes. |