Majestic Labs Unveils Prometheus AI Server to Challenge Nvidia GPUs

Founded in 2023 by former Google and Meta engineers, Tel Aviv-based startup Majestic Labs unveiled the Prometheus server. The system replaces costly Nvidia GPUs and high-bandwidth memory with custom Arm-powered Ignite AI Processing Units and up to 128TB of unified LPDDR6 RAM, targeting the fundamental memory-bound constraints of AI inference.

The artificial intelligence industry has hit a thermal and economic wall. For years, scaling large language models meant chaining together massive arrays of graphics processors and expensive high-bandwidth memory stacks. It is an architecture that delivers blistering compute speeds, but chokes on memory capacity and bandwidth when deploying massive models.

Smashed Memory Walls and Custom Silicon

Majestic Labs is taking a fundamentally different route to scale. Instead of placing memory directly on GPU packages, the Prometheus server deploys up to 12 Ignite AI Processing Units (AIUs). These processors combine Arm cores with RISC-V vector and tensor engines, pooling between 8TB and 128TB of contiguous, coherent LPDDR6 memory.

How does that memory stay connected without creating a massive latency bottleneck? Custom memory aggregation chiplets handle the heavy lifting. These chiplets link together using copper cables up to one meter long. A standard 40U rack accommodates four Prometheus servers, drawing a total of 120 kilowatts and relying on cold-plate liquid cooling systems rather than forced air.

Let us look at the hardware disparity. An Nvidia DGX B300 system outfitted with eight Blackwell GPUs provides 2.3TB of HBM3e alongside a ceiling of 4TB of standard DDR5 system memory. Majestic Labs claims its architecture delivers over 50 times more fast memory than that rival configuration, coupled with 1.7 times its interconnect bandwidth.

“One Majestic rack holds the fast memory capacity of 25 Nvidia NVL72 Vera Rubin racks at a fraction of the power,” Majestic Labs stated regarding the system’s efficiency.

The Architecture and Software Stack

The enterprise infrastructure play here is not just about raw density; it is about lowering the barrier to entry for workloads that previously required hyperscaler-tier budgets. Majestic Labs asserts that organizations priced out of traditional GPU clusters can now run demanding AI inference workloads locally. In theory, this yields up to 1,000 times more memory available per processor.

Under the hood, the hardware is designed to be Open Compute Project (OCP) compliant. More importantly for developers, it supports PyTorch, vLLM, and OpenAI’s Triton frameworks. This ensures that existing model weights and inference pipelines can run without requiring custom code rewrites or proprietary driver adaptations.

Prometheus Server vs. Traditional GPU Clusters:

  • Processors: Ignite AIUs featuring Arm cores and RISC-V vector/tensor engines.
  • Memory Pool: 8TB to 128TB of unified LPDDR6 RAM per server via memory aggregation chiplets.
  • Cooling & Form Factor: Cold-plate liquid cooling in a standard 40U rack holding four servers.
  • Ecosystem Support: Native compatibility with PyTorch, vLLM, and OpenAI Triton frameworks.

Market Realities and Unanswered Questions

Founded by CEO Ofer Shacham, President Sha Rabii, and COO Masumi Reynders, Majestic Labs employs roughly 40 people across Tel Aviv and Los Angeles. The company secured a $100 million Series A funding round in late 2025 and reports significant initial orders from large enterprises, neoclouds, and hyperscalers ahead of a planned shipping window next year.

The startup projects that Prometheus servers could cost between 10 and 50 times less than an equivalent GPU-based system while consuming significantly less electricity per rack. Yet, independent validation remains entirely absent.

Hardware analysts have pointed out glaring logistical questions regarding the 128TB configuration. Building a single server with 128TB of LPDDR6 memory using widely available 2GB LPDDR6 dies requires approximately 64,000 individual memory dies. That scale implies well over a hundred memory aggregation chiplets packed into a single server chassis.

Enterprise buyers evaluating a shift away from established graphics processor ecosystems will need to review independent benchmarks and thermal telemetry once silicon actually lands in data centers.

The 30-Second Verdict

Majestic Labs is betting that the future of AI inference belongs to massive, unified memory pools rather than isolated, high-bandwidth GPU caches. If the company delivers on its shipping schedule and pricing projections, it could disrupt the hardware monopoly currently enjoyed by traditional accelerator vendors. Until independent testing validates the real-world performance of its Arm and RISC-V chiplet architecture, however, these striking numbers remain strictly on paper.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

Denis Shapovalov Defeats Cameron Norrie in Los Cabos Open Semi-Final

Germany Deported Bob Vylan Frontman Over Pro-Palestinian Comments

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.