Photo by Vitaly Gariev on Unsplash. Source: https://unsplash.com/photos/man-using-laptop-on-sofa-AYAuwXLE7Ag (Unsplash License).

Executive Summary

Ghost, a San Francisco startup founded by 19 year old Zain Javaid, left stealth on October 5 with an $11 million seed round led by Andreessen Horowitz and a $3,499 computer called Core. The box has no screen. It runs open weight models on a bundled Nvidia RTX Pro 4000 SFF Blackwell GPU and keeps both the data and the model weights on the device.

The finding is that local AI inference crossed from hobbyist project to a product people pay for. The first batch sold out within hours. That is a rounding error next to the hundred billion dollar cloud buildout, but it is the first real price set on the other side of the bet, that a slice of users want a model they own with no per token bill and nothing leaving the house. Cloud inference is still where the volume is. It is no longer the only option.

Every dollar in this year’s AI buildout rests on one assumption. That the model runs somewhere else, answers to a rate limit, and bills by the token. Ghost just sold its first batch of the opposite.

For $3,499 the company ships Core, a screenless computer that runs agents locally. Users drive it from a phone app. Under the lid sits a six core AMD Ryzen 5 7600, 64GB of DDR5 memory, a 1TB NVMe drive, and an Nvidia RTX Pro 4000 SFF Blackwell GPU with 24GB of its own memory. The GPU is the point. It is bundled into the price, so a buyer does not source a card separately.

Four open weight models ship preinstalled. Qwen 3.8 Next, Qwen 3.8 27B, Gemma 4 31B and Muse Glimmer 30B. A user can load others from Hugging Face or bring their own weights. Ghost says a firewall watches every outgoing request and blocks anything that looks like data leaving the device without permission.

Diagram. Local inference compared with cloud inference. The Ghost Core box runs open weight models on a bundled Nvidia RTX Pro 4000 SFF Blackwell GPU and keeps data on the device.
Cloud inference bills by the token. The Core box is bought once.

The sellout is the signal, not the spec sheet

Preorders opened on October 5 and the first batch sold out within hours. Founder Zain Javaid wrote on X that Ghost locked in its GPU supply at about $1,700 and that the card has since doubled, so the next batch will likely cost more. A one time price that rises is still not a per token meter.

Put the arithmetic next to a subscription. At $20 a month, $3,499 is about fourteen years of a chatbot plan. The difference is what you are buying. A subscription rents access to a frontier model. Core buys a fixed box that runs 27 to 31 billion parameter open models and keeps working if the vendor disappears.

The constraint moves to the edge, and the controls follow

A 24GB card sets a hard ceiling. Core runs capable open models. It does not run the largest frontier systems, which sit behind APIs because they need clusters no shelf can hold. So this is not cloud inference replaced. It is cloud inference unbundled at the low end, where the workload is personal, private and latency sensitive.

For enterprises the same physics applies one rack up. A regulated team that cannot send prompts to a third party buys the GPU and owns the data path. That is the argument behind on premises inference on Kubernetes, and it is why the interesting part of Core is not the chip but the supervision layer around the agent.

What a small startup has proved, and what it still has to prove

The bet is legible now. Andreessen Horowitz led an $11 million seed into a small team building units by hand in San Francisco. That is a tiny check against the capital moving into AI compute. It is also a live test of whether buyers will pay upfront for a model they own.

The open questions are support and shelf life. A small company has to ship driver updates, model updates and a warranty across a device meant to outlast the model inside it. Ghost offers a one year warranty and a 30 day return. Buyers are betting the company stays alive past the novelty.

Three questions to apply to your own environment. Which inference are you paying for by the token that you could own outright? What data would you keep on device if the hardware made it easy? And if the price of the accelerator keeps climbing, what does that do to the one time buy?

Related reading. Model API pricing now spans a 42x range, which is the bill local inference competes against. A sandbox escape at Vercel showed why agent isolation matters. And Nvidia moving its agent guard off the host is the enterprise version of the same control.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.