It looks after itself

The best outage is the one that never happened.

A node starts failing hours before it stops. Genie sees it, moves the work off while nothing is being interrupted, and you find out afterwards, if at all.

Failing over fast is the easy half

When a node stops, the cluster fences it and promotes a replacement in about twelve seconds, in code, with nobody woken. That part is solved and it is the part everybody talks about.

What nobody does is act on the hours before it. A disk crossing its own growth line. Memory that stops coming back between cycles. An error rate rising on one node and no other. A scheduler lease acquired and lost, over and over. All of it is measured, and on every other platform all of it just sits there until something breaks.

Genie drains the node instead. Sessions finish where they started, collection moves to its neighbours, the lease is not renewed, and the node comes out of rotation with nothing interrupted. A clean drain and an emergency failover reach the same end state and feel completely different to your client. One is invisible. The other is a report that did not run.

The same failure, twice · 1:48

And when it recovers, it comes back the same way.

Why an analyst beats a threshold here

A rule fires on one fact. A disk at ninety percent. A queue over a hundred. Every platform in this market has thousands of them and every MSP has learned to ignore most of them.

Genie holds all the facts at once and knows which combination means something. A disk climbing while a retention prune has quietly stopped succeeding is a finding. The same disk climbing after a restore is not. A rule that could tell those apart would have to be written by somebody who had already seen both.

And nobody is woken for something he fixed. A service restarted, a job re-scheduled, a maintenance task re-run, space reclaimed within the policy you already set — recorded, not announced. An operator notified about fifty resolved things stops reading the fifty-first.

What does reach you is what he could not fix, and it arrives as his account of it — what happened, what he tried, whether it will recur, and what he could not determine. Not a raw alert at 2am for somebody to reconstruct from a fragment.

And a thing healed three times is not healed. A repeating repair is escalated as a pattern, with its interval and what he believes the cause is — because the fix he keeps applying is the symptom.

A fault found on one deployment is fixed on all of them

Every customer runs the same software on their own hardware. When something goes wrong on one box, every other box has the same code and nothing checks whether they have the same problem.

Here they do. The health record you already send carries what happened, anonymously and signed. A pattern across deployments becomes a finding, the finding becomes a fix, and the fix reaches everybody — usually before the practices carrying it noticed anything.

And you can read the whole loop. Every release names what it closed, in the same open-items list you already have access to, and where the finding came from the fleet it says so. A fault found on somebody else’s deployment, fixed before it reached yours, and written down.

Nothing ships from a pattern alone. A finding becomes a fix the same way everything else does — specified, tested against the case that produced it, and proven before it goes anywhere. The loop is fast. It is not unsupervised.

Your standby engine takes every release first and has to prove it came back before the primary is touched. A release that behaves worse than the one before it halts on its own. And an air-gapped deployment gets every fix by the same file, losing nothing but the automation.

And it keeps getting quieter

Your card is idle most of the day. A practice asks a few dozen questions and the hardware sits there for the other twenty-three hours.

So Genie keeps working. Going back over findings he raised months ago against everything learned since. Re-examining machines that were checked when less was known. Comparing what your documentation says against what is actually there. And reviewing his own wrong answers, which nothing else in this market does at all.

A false positive gets retired rather than tuned around. Which means the noise falls month over month instead of staying flat — and that is a difference you can measure in your own ticket queue.

It expands into whatever the card is not using and stops the instant you ask it something. A technician never waits behind a thought nobody asked for, and there is always headroom left for the morning everything happens at once.

And none of it costs you anything, because the card is already yours and the electricity is the same whether he thinks or not.

How many devices one box carries

Collection is what caps an estate, not the AI. A practice with five thousand devices asks a handful of questions a day and polls five thousand devices every cycle.

One box does 3,575 polls a second. That is measured, on a single machine, with sixteen collectors running — and it was still getting faster when our test rig ran out of load to give it. So it is a floor, not a ceiling, and we say so rather than rounding it up.

What that is in devices depends on how hard you check them, and the honest answer is a range. On a lightly monitored estate it divides out at about 53,000 devices. On a heavily instrumented one, about 12,500. Both are measured, and both are floors.

A five thousand device practice is a large MSP, and the smaller of those two figures is more than twice that. Most practices will never need a second box for capacity.

The second one exists for a different reason — a subnet only one machine can reach, or a machine that stops. Not throughput.

And we will tell you your number rather than ours. Bring your device count and how often you want each one checked, and the arithmetic takes a minute. It is the same arithmetic we used, and you can check it.

So the collectors scale sideways inside one platform — not another licensed server per poller the way this industry has always done it. Devices are distributed automatically, a collector that stops has its work picked up, and you are never asked to reconfigure anything.

And most of what a monitoring platform collects has not changed. A switch’s counters move every second. Its chassis inventory has not changed in three years and gets walked every five minutes anyway. Genie stops asking questions whose answers never move, and starts again the moment anything suggests he should.

Availability never adapts. Whether a device is up is cheap and it is checked at the rate you set, forever. What adapts is the expensive detail — and anything you mark critical is exempt entirely, no reasoning and no exceptions.

That is a lever no rules engine has, because telling a stable printer from a quiet failure takes an analyst that has watched that specific device on that specific network for months.

Try it on your own estate

Thirty days, read-only, on your own hardware. No card, no call, and nothing to uninstall if you walk away.