We size the box with you Tier 1 is measured on the card it names. Tiers 2 and 3 are derived from that measurement and say so in the table, because every other vendor prints a sizing table as though every row were a specification. Tell us your asset count before you buy anything and we will size it with you — that conversation is free, it takes ten minutes, and it is the reason nobody buys the wrong card twice. replaces the estimate.

Documentation  ·  Getting started

One box for your whole practice.

Not one per client. Every client you manage lives inside a single deployment, isolated by tenancy rather than by separate installs. A fourteen-client practice runs on one server.

The host

CPU8 cores minimum
RAM32 GB minimum
Disk500 GB SSD, grows with retention
OSLinux with Docker
Networkreachable from your sites

Disk is the one that surprises people. Log and metric retention drives it, and long retention on a busy estate is measured in terabytes, not gigabytes.

UP TO 5,000 ASSETSRTX 4000 Ada ×120 GB VRAMMEASURED5,000 - 10,000RTX 4000 Ada ×240 GB VRAMDERIVED10,000 - 20,000RTX 4000 Ada ×480 GB VRAMDERIVED20,000 +RTX 4000 Ada ×8160 GB VRAMDERIVEDRAM and disk are driven by retention, not by the card — the host table above is the floor.16GB cards are disqualified.The reason is measured rather than asserted — the model does not fit.

The GPU

We measured one card and published the whole run. Everything else on this page is arithmetic from published specifications, labelled as such in the cell rather than in a footnote.

12 GB — runs, with no headroom. This is the floor.

Measured on an RTX 4080 Laptop: a 14B model at an 8,192-token context occupies 11.5 GB of 12 GB, leaving about 700 MB free, and serves eight concurrent requests at 38–42 tokens per second each. It works. What you do not get is room — ask for a context window this card cannot hold and throughput falls by a factor of 2.5 as the weights spill to host memory.

16 GB is the first size with headroom. 12 GB is the first size that works.

Card VRAM Bandwidth Generation speed
RTX 4080 Laptop 12 GB 432 GB/s 39.5 tok/s — measured
RTX 4060 Ti 16GB 16 GB 288 GB/s 26 tok/s (projected)
RTX 3090 (used) 24 GB 936 GB/s 85 tok/s (projected)
RTX 4090 24 GB 1008 GB/s 92 tok/s (projected)
RTX 5090 32 GB 1792 GB/s 164 tok/s (projected)

The method, stated so you can argue with it. Single-stream generation of a dense model is memory-bandwidth bound — every token read reads the whole weight set once — so throughput scales with bandwidth for the same model. The projection is our measured figure multiplied by the ratio of NVIDIA's published bandwidths. Nobody has run this product on those cards.

What the method does not cover, and it is not small. Time to first token is prefill, which is compute-bound rather than bandwidth-bound, so there is no column for it — a blank is better than a number nobody can stand behind. Concurrency depends on how much VRAM is left for the cache after the weights, which differs on every card.

For inference, VRAM matters more than speed. The deciding question is whether the model and its full context fit. A card that fits does not pay the 2.5× penalty; a faster card that does not fit still pays it. That is the single most useful sentence on this page for choosing a card.

Many MSPs already have a virtualisation host with room for a card. Ask us before buying anything.

Which card, specifically

You do not need a datacentre GPU and we will not pretend you do. This is a consumer card in a machine you probably already have room for.

Card VRAM Typical
RTX 3090 — used
Our pick. Best VRAM per dollar on the market.
24 GB $700–1,000 USD
RTX 4060 Ti 16GB — new
The floor. Runs it, with warranty, at 165 W.
16 GB ~$450 USD
RTX 4090
the same 24 GB as a 3090. Same VRAM.
24 GB $1,600–2,000 USD
RTX 5090
More than you need. Priced accordingly right now.
32 GB $3,300–4,600

GPU prices are unusually high in 2026. A memory shortage has pushed new cards well above list, and a 5090 that launched at $1,999 sells for more than twice that.

Which is exactly why we point at a used 3090. For inference, VRAM matters more than speed — and 24 GB for under a thousand dollars is not close to being beaten. Inference also does not stress a card the way gaming or mining does, so a used one is a reasonable buy.

You do not need an A6000, an L40S, an H100, or anything with a datacentre badge on it.

Those cards exist for training and for serving thousands of concurrent users. You are running one model for one practice. A five-figure card would sit idle and you would have paid for the privilege.

Prices are what cards actually sell for in August 2026 and they move weekly. Many MSPs already have a virtualisation host with a free slot — ask us before buying anything.

On a 4090 or a 3090 we quote arithmetic, not a benchmark. Memory bandwidth alone puts it around 8% ahead — that figure comes off the specification sheets, and we say so rather than dress it up as a result.

At each client site

Collector2 cores, 4 GB, 20 GB disk
Spool10 GB store-and-forward
Endpoint agentnegligible
Outboundto your deployment only

The spool is why a dropped WAN link is a delay rather than a gap. When the link returns, the collector backfills — and a collector that stops reporting raises an alert, because a silently unmonitored client is the failure this product exists to end.

Air-gapped

Fully supported, and not a lesser mode. The platform makes no outbound call of any kind in normal operation — the model runs locally, updates are files you carry in, and the licence is a signed file you download and copy across.

The one honest cost: an air-gapped deployment receives licence revocations and product advisories later than a connected one, because it receives them when you next carry a licence in. Documented rather than hidden.

How many clients will a box hold?

The honest answer is that it depends on your estate, and we would rather size it with you than sell you a number off a table. What follows is how we work it out, so you can do the arithmetic yourself.

Every other vendor's sizing table is a guess dressed as a specification. We tell you which figures are arithmetic and which are measured, rather than publish a number you would plan a purchase around and then discover was optimistic.

What we will publish, once the load test runs: client companies, managed assets, events per second and concurrent Genie operations for each hardware tier — with the recorded run behind it, and the point at which the box degrades.

A growing practice adds a second collector rather than a bigger box, and the ceiling we publish is measured rather than projected. How collection scales, and what breaks first.

If a tier's real ceiling turns out to be lower than we hoped, the documentation says the lower number. That is how 5,000 got published rather than a rounder figure somebody liked better.

Tier 1 carries 5,000 assets, and that figure is measured. It comes off a load run of 2,905 inferences an hour on an RTX 4000 Ada with 20 GB, not off a specification sheet.

Tiers 2 and 3 are derived from that measurement and labelled derived in the table, in the cell rather than in a footnote. We publish a client count when a production deployment gives us operations per client per hour — assets we can measure, client companies we cannot, and we would rather give you the number we have than multiply it by a guess.

What degradation looks like

When a deployment is under-provisioned it does not fail silently, and it does not quietly get worse at its job.

Work queues, it does not vanish

Background analysis waits. Interactive work and anything safety-critical keep their capacity.

Genie says he is slow

He tells you the deployment is saturated. He does not silently return a shallower answer and let you believe it was his best.

Monitoring never sheds

Collection, alerting and detection are the last things to give up capacity, and they are never dropped to make room.

It tells you before it hurts

Capacity headroom is reported continuously, with an estimate of when your growth will exhaust it.

Sizing is the question we most often get asked before anyone has run anything. Describe your estate and we will tell you what we actually know and what we do not — including if the honest answer is that you should wait for the measured numbers. hello@dream-genie.ai

Already have a machine that fits?

Start a thirty-day evaluation and find out what your current tools are hiding.

Growing, and surviving a node

Most practices run one machine and that is a supported answer — if it stops, your agents keep queueing and nothing is lost. For a practice that cannot accept that, the platform runs across three nodes with quorum, fencing before promotion, and one model of the truth across every one of them.

Two nodes is not high availability and we will not call it that. Two cannot establish quorum, and selling a pair as HA is discovered during the first real failure.

And growing is additive rather than a migration. A second GPU serves more requests at once; a third node means surviving the loss of one. No client is ever moved onto a different box, because no client lives on one.

The topology, the failover sequence, how a deployment grows and what each shape does not survive: high availability.

Try it on your own estate

Thirty days, read-only, on your own hardware. No card, no call, and nothing to uninstall if you walk away.