A small box instead of a server in the living room
What interests me about this new wave of AI boxes is a straightforward promise. A compact machine that fits in an office and runs models locally, without installing a server weighing tens of kilograms. Inference, meaning running the model, becomes something you can buy and manage on your own premises.
These products serve different needs. Framework offers a hardware platform, Lucebox a workstation with its inference engine preinstalled, and Ghost a computer built around a personal assistant. Software readiness matters as much as memory capacity. An assembled computer does not necessarily arrive with a business agent ready to work.
I have not tested these three boxes. This article compares their announcements as of October 6, 2026 with my experience on my own PC. Store screenshots document prices displayed on that date, rather than immediate availability or measured performance.



I previously explored the Xiaomi AI Cube, the “Xiaomi box” positioned against DGX Spark, described as a prototype in that article, and Apple's Mac Studio for local LLMs. Framework, Lucebox and Ghost broaden that landscape with different architectures and levels of software readiness. Prototypes, preorders and shipping products need to remain clearly distinguished.
What stands out to me is how quickly the options are multiplying. The choice extends beyond NVIDIA versus Apple, with large unified-memory workstations, hybrid systems with discrete GPUs and integrated assistants. That momentum makes software, support and actual workloads more important than a chip name alone.
Framework, Lucebox and Ghost take three different approaches
| Machine | Announced memory and compute | Observed price | Announced availability |
|---|---|---|---|
| Framework Desktop AI Max 400 | Ryzen AI Max+ PRO 495, 192 GB unified memory, integrated GPU | €7,659 in the supplied French configurator screenshot | Preorder, displayed batch scheduled for November 2026 |
| Lucebox Zero 495 | 128 GB unified memory and 32 GB Radeon AI PRO R9700 ; optional 192 GB unified memory | From €6,499 ; €8,698 with the 192 GB option, before other upgrades | Shipping announced for January 2027 |
| Ghost Core | Ryzen 5 7600, 64 GB DDR5 and 24 GB RTX PRO 4000 Blackwell SFF | $3,499 displayed, French delivery and tax treatment to confirm | First batch open for orders ; delivery date not confirmed here |
The 192 GB Framework Desktop, announced on September 30, adds capacity with 8,533 MT/s LPDDR5X and a stated 273 GB/s bandwidth. Its Fedora prebuilt and DIY editions offer different levels of convenience. This generation is designed for Linux, and the configurator lists the memory as non-upgradeable. Its 192 GB is shared with the system, rather than entirely reserved for model weights.
The Lucebox Zero 495 combines unified memory with a discrete GPU. Upgrading from the base 128 GB to 192 GB adds €2,199. The seller lists taxes, duties and shipping as included in its euro price, with Ubuntu Server and its engine preinstalled. Its “224 GB” figure adds 192 GB to 32 GB of VRAM. That is not a homogeneous pool operating at GPU speed throughout, and the PCIe 4.0 x4 connection brings its own constraints.
Ghost Core has separate 64 GB system RAM and 24 GB VRAM. It is positioned around a personal local assistant connected to applications and files. Listed models include Qwen 3.8-27B and Qwen 3.8-Next, but those names alone do not establish throughput, quantization or supported concurrent sessions.
Which machine fits which model and workload?
Here, B means one billion parameters. These are my sizing estimates, not manufacturer-certified maximums or measured performance. I assume roughly 4-bit quantization, budgeting 0.6 GB per billion parameters for weights and their metadata. Context, the runtime and the operating system require additional memory.
| Configuration | Dense models to consider | MoE and estimated Q4 ceiling | Where I would use it |
|---|---|---|---|
| Framework Desktop 192 GB | Start with 32 to 70B ; consider 120B if throughput is acceptable. Estimated memory ceiling of 260B, without a speed guarantee. | Up to 260B total parameters, depending on runtime support. A 125B MoE leaves more headroom than a model near the ceiling. | Document RAG, development and experimentation with large models on Linux, with control over the software stack. |
| Lucebox Zero 495, 128 GB + 32 GB GPU | 32 to 40B can target the GPU alone ; 70 to 120B require splitting weights. Aggregate ceiling of roughly 200B if supported by the runtime. | Up to 200B total, split across memory tiers. Frequently used experts may benefit from the GPU, depending on the engine. | Agents and RAG for a small team, local inference services and MoE with preinstalled software. Concurrency still needs testing. |
| Lucebox Zero 495, 192 GB + 32 GB GPU | Same GPU-only capacity, more room for split weights. Aggregate ceiling of roughly 300B, potentially slow for dense inference. | Up to 300B total under these assumptions. More memory increases capacity, not automatically throughput. | Large MoE models and experimentation. Leave room for context and documents instead of always loading the largest model. |
| Ghost Core, 64 GB RAM + 24 GB GPU | Start with 7 to 27B ; roughly 30B on the GPU alone depending on format and context. A dense 70B needs RAM offloading and may be slow. | Aggregate ceiling of roughly 110B in Q4, with a runtime supporting that split. This calculation does not establish support in Ghost's supplied software. | Personal assistance, email, filing, file search and everyday automation, prioritizing convenience. |
| My PC, 24 GB RTX 3090 + 64 GB RAM | Similar memory envelope to Ghost, not an assertion of equal speed. Start with 7 to 27B, or 70B with offloading. | Roughly 110B in Q4 by this calculation. My actual test uses 125B in the more compact IQ3_XXS format, with Strata and around 70 tokens/s at a configured 128K window. | Developers or enthusiasts reusing their hardware for agents and RAG, accepting setup, tuning and maintenance. |
Total MoE parameters matter for weight storage, while active parameters mainly affect compute. A 125B MoE activating only a fraction of its experts does not have the memory footprint of a dense model with that active parameter count. Hugging Face's MoE explanation describes this distinction. At comparable quantization, dense and MoE weight budgets are of a similar order per total parameter, but throughput and context cache requirements can differ substantially.
For reproducibility, I reserve 32 GB on unified-memory configurations, 16 GB on 64 GB system-RAM machines, and 8 GB or 6 GB on 32 GB or 24 GB GPUs respectively. That leaves 160, 120, 184 and 66 GB for weights across the four distinct configurations, which I divide by 0.6 and round down. Adding RAM and VRAM assumes a runtime that can split weights without excessive duplication. It never means a single pool operating at GPU speed.
Those reserves do not guarantee a 128K or 262K context. Long contexts, concurrent users or some KV caches can need much more. Q8 and BF16 take more space ; Q3 or Q2 may exceed these ceilings with quality tradeoffs to evaluate. SSD offloading can push loading limits further, but I do not count it as usable capacity here. Before buying, validate the exact weight file, runtime, context and throughput on a representative task.



Price per gigabyte alone is not a meaningful ranking. The comparison needs to include supplied software, maintenance, storage, warranties and the complete configuration with consistent tax treatment. Several thousand euros remains a substantial investment for an independent professional.
DGX Spark still needs an accurate comparison
On October 2, NVIDIA announced a 64 GB DGX Spark variant, starting at $4,999 through partners and scheduled for October 23. At that price, I find the memory capacity difficult to justify when loading large models is the main goal. Its software environment may nevertheless matter to a professional buyer.
This is different from the 128 GB DGX Spark in my screenshot. That listing shows an out-of-stock product without a visible price. The two configurations should not be confused or compared directly with European tax-inclusive prices without adjustment.

These offers extend the discussion in my DGX Spark and GB10 machine comparison. But avoiding a purchase is becoming an equally interesting part of the story.
My PC already runs Qwen with Strata
I bought my RTX 3090 used on Le Bon Coin for €750. I also have 64 GB of RAM bought for a few hundred euros several months ago. Those are my historical purchase costs, not a reproducible quote today or the price of the complete PC.
With Strata, I use Qwen 3.8 Flash with IQ3_XXS quantization, which reduces the storage required for model weights. The project supports Qwen3.8-Flash-Next, a 125-billion-parameter model. I did not record the exact engine version for this report.
My reported throughput is around 70 tokens/s with a 128K window, and around 30 tokens/s with a 262K window. These are observations from my use, not a controlled comparison with identical prompts and settings. Tokens are the units of text produced by the model. The October 5 screenshot displays 73.8 tokens/s during generation, with about 46K of the configured 128K context occupied, or 35%. It therefore does not demonstrate that speed with a completely full 128K context.

The monitor separately shows prompt processing at 1,745 tokens/s. It reports 23.8 GB of 24 GB VRAM in use, around 60.5 GB of 64 GB system RAM, and instantaneous GPU power of 332 W. Its history includes a completed request at 36.2 tokens/s and an error whose cause is not visible. I show the complete screenshot because it documents a real session and its variations rather than an isolated laboratory result.
Even with those limits, this is already a credible alternative for my agent workflows, in which the model chains tool calls. I can choose the model and engine, keep inference on my machine and decide which permissions to grant. That gives me practical control without buying a new workstation.
Software is pushing memory limits further
Mixture-of-experts models, or MoE, activate only some experts for each token. Strata’s technical documentation describes retaining experts in VRAM, using system RAM and reading weights from files when memory is insufficient. This placement avoids requiring all weights to remain on the GPU.
An expert cache is different from the KV cache, which stores information associated with the conversation context. Moving weights to an SSD does not turn it into VRAM. Transfers, bandwidth and latency still matter, particularly as context grows. Fitting a model and serving it comfortably are different objectives.
TensorFold also targets efficiency, including speculative decoding that proposes and verifies several tokens. It supports Apple Silicon and NVIDIA. At the time of consultation, CUDA kernels require compute capability 8.9 or newer, covering RTX 40 and compatible newer hardware, while RTX 30 cards are explicitly rejected. The project also notes incomplete hardware validation on some cards. I am therefore making no claim of current RTX 3090 compatibility or a future support date.
These projects do not make every model work on every computer. They do show that buying more memory or a new GPU is not the only way forward. The engine, weight format, context length and request handling can substantially change the result.
Use the strongest model where it actually helps
Calling the largest model to rename a file or classify a message can feel like using a bazooka to kill a fly. Some operations do not need AI at all. For others, a suitable local model may be enough, while a frontier model adds more value for difficult scoping, complex planning or demanding creative work.
No model
Rename, move or filter using explicit rules when no interpretation is needed.
Suitable local model
Classify, draft or retrieve from documents, with constrained tools and permissions.
Frontier model
Complex planning or demanding work, only when data and policy permit external processing.
This follows the approach in my article on self-hosted Hermes and sensitive data. A local agent can draft emails, classify documents or support RAG, retrieving information from a document collection to ground its answers. Sending, deleting, booking and sensitive decisions still require validation. The harness, meaning the surrounding tools, permissions, checks and recovery mechanisms, is part of the result’s quality.
A rate of 20 to 50 tokens/s can feel comfortable for many of these tasks. It does not predict the duration of an entire workflow involving document ingestion, search and tool calls. A fast answer does not demonstrate accuracy either.
The idea is credible for a few people using the machine intermittently. Several simultaneous generations with long contexts require measurements of waiting time and memory consumption. A cluster becomes necessary because of workload and availability requirements, not simply because an organisation has reached a certain size.
The efficiency race is also changing chips
Desktop boxes are only part of this shift. Google's Ironwood TPUs target large-scale AI, with a generation designed particularly for inference. OpenAI has also introduced Jalapeño with Broadcom and Celestica, a chip tailored to LLM workloads, with initial deployment announced for late 2026. These are infrastructure strategies, not add-in cards available for our PCs.
Anthropic reportedly discussed chips with Samsung. No confirmed agreement here.
Another approach moves compute closer to data. XCENA's MX1 is a computational-memory platform using CXL, an interconnect linking processors, memory and accelerators. It combines DDR5 with processing cores and provides for SSD-backed expansion. Near-data processing aims to reduce transfers instead of merely adding central compute power.
XCENA's software tools target workloads including analytics executed close to memory. That does not turn an SSD into a universal GPU or demonstrate that an entire LLM runs faster than on the machines compared here. It addresses data movement and capacity through a different architecture.
To me, these approaches point towards the same shift. Efficiency depends on models, inference engines, memory and increasingly specialized chips. Some innovations initially target data centers, while others reach our desks. More approaches mean more possibilities, without guaranteeing that each will quickly become an affordable local product.
Size the investment around a need, not a fear of missing out
Memory-market pressure is documented. TrendForce’s September 30 outlook forecasts quarterly contract-price increases of 10–15% for conventional DRAM and 15–20% for NAND in the fourth quarter. These are component-market forecasts, not a guaranteed uniform increase for every SSD, GPU or computer.
I understand the appeal of securing a useful configuration at a known price. A launch offer nevertheless guarantees neither the best deal nor immediate delivery. Before spending €6,000 to €9,000, I would test the intended model on existing hardware, then estimate actual cloud savings and time saved.
Electricity, administration, backups and useful lifetime all belong in that calculation. Simply spreading €6,499 over three years works out at about €181 per month for the hardware alone, before those expenses. Light API usage may be cheaper. Regular workloads involving data you want to keep on premises can justify dedicated hardware.
For a small practice, business or independent professional, privacy may matter as much as the financial calculation. Local inference reduces exposure to an external model service, but an email connector, web search or telemetry can still communicate outside the machine. Access, data flows and updates still need control. Hardware ownership provides autonomy, not automatic compliance.
My bet on the next gold rush for efficiency
This development extends beyond desktop boxes. Google already offers Gemma 4 E2B and E4B through its Android AICore developer preview on compatible devices. That illustrates the growing role of models matched to hardware rather than systematically calling remote infrastructure.
To me, the next gold rush will be the race for efficiency. Progress in hardware and software suggests that laboratories and industrial players are converging on the same goal, doing more with the resources available. As AI demand accelerates, constraints on data-center grid connections, accelerator supply and memory availability make that pursuit unavoidable in my view. We have moved beyond a proof of concept and into growing everyday use. I do not expect us to turn back, but to change how we consume AI.
My scenario is a gradual easing of pressure on GPUs as alternative compute platforms and better software take on some workloads. I am less convinced that RAM and SSD constraints will ease, since these architectures still need them. Even so, I can envisage some hardware and rental prices stabilizing or falling slightly over the coming months. That is my expectation, not an assured or across-the-board price decline.
I do not expect API prices to rise indefinitely either. However, an unchanged bill can conceal fewer included tokens or tighter quotas. That would increase the cost of the same usage without changing the advertised price. Technical efficiency therefore does not guarantee that customers will immediately receive all the savings.
I am not betting on AI suddenly coming to a halt. Competition, including Chinese open-weight models and laboratories worldwide, looks more likely to broaden access. We will distribute more workloads across computers, phones, local boxes and the cloud, using each resource where it adds value. I believe AI will become increasingly present through a hybrid approach, and that the best is still ahead.
Sources and methodology
Sources consulted on October 6, 2026. Hardware specifications and schedules are vendor announcements rather than independent measurements. My Strata experience is a personal report without a controlled comparative protocol.
Framework’s September 30 announcement and French configurator. The supplied screenshot documents the displayed €7,659 price for 192 GB ; the research tool could not retrieve the configurator itself.
Lucebox store and configurations and engine overview. Pricing, memory upgrade, architecture and announced shipping schedule.
Ghost Core official page. Dollar price, CPU, RAM, GPU and software positioning.
NVIDIA’s October 2 DGX Spark 64 GB announcement and 128 GB product page. The 64 GB announcement and price were found through indexed official search results ; the full announcement could not be retrieved during this check.
Strata repository and technical documentation. Supported model and expert-memory management. My screenshot’s throughput is not a project-provided benchmark.
Google Ironwood, 2025-04-09.
OpenAI Jalapeño, 2026-06.
The Information, 2026-07-02.
Hugging Face, Mixture of Experts Explained. Total and active parameters. Capacity ceilings are my estimates, not vendor benchmarks.
TensorFold repository documentation. CUDA compatibility and limitations observed on this date, which may change.
TrendForce’s September 30 forecast. Expected changes in DRAM and NAND contract prices.
Google’s April 2 AICore developer preview. Local models on compatible Android devices, found through indexed official search results.
Product and pricing screenshots are dated October 6 ; the Strata screenshot is dated October 5. French storefronts are shown as captured. The AI-generated cover uses the supplied enclosure screenshots and an RTX 3090 Founders Edition reference. It is not a photograph of hardware I have personally tested together.



