Two engines · one small on purpose
The thing that watches must not die with the thing it watches.
A small engine beside the platform. It maintains the box overnight, it survives what stops everything else, and it can do four things and no more.
And what the pair of them do when something starts going wrong.
What survives what, and what each engine may do
Forty years old, and new here
Every server you own has one. iDRAC, iLO, IPMI — a small processor that stays up when the operating system does not, so somebody can find out why. Firewalls separate the control plane from the data plane for the same reason. The principle is that the thing that watches must not share fate with the thing it watches.
No MSP platform has one, because no MSP platform runs on your hardware. There is nothing to supervise when the product is somebody else's cloud. This one runs on your box, so the question becomes real.
What it can do, in full
Read health. Restart a declared service. Run a declared maintenance job. Report to us.
That is the complete list, and it is shorter than the agent's on purpose. Every out-of-band engine ever shipped became an attack surface for the same reason: a privileged process that survives when everything else is down is exactly what somebody would want. The safeguard is that it is small, not that it is trusted.
It never reaches a client system. Not restricted, not audited, not permission-gated — absent. There is no route, no credential and no path. The agent talks to your estate; the service engine talks to the box it is on, and to us.
What it does while you are asleep
Vacuum, reindex, retention pruning, certificate checks, and restoring a backup to see whether it restores. All of that runs anyway. What is different is that something reads the result.
A vacuum ran every night for six weeks and the table grew fourteen percent. The job is green. The result is not. A scheduler cannot tell those apart, and that is the difference between a cron entry and an analyst.
When the platform stops
And by the time you are running two nodes and two collectors, that is five things it survives, not one. The platform serving, the platform collecting, and the devices being polled all go dark together. The small box on the right does not.
The service engine is still up, and it is the only thing still talking. What it sends is what it observed, what it tried, and what it could not determine — and we call you, rather than waiting for you to notice.
It may restart a declared service. It may not repair, reconfigure, migrate or reinstall. When restarting does not work it says so and stops, and a person here is told.
And the other way round changes nothing. If the service engine stops, your platform serves, collects and answers exactly as before. You just stop hearing from us about the box. The supervisor's failure is our problem; the platform's is yours, and the supervisor exists to make yours shorter.
Neither engine updates itself
Each updates the other. The one that stays behind is the one that fixes the one in front. iDRAC firmware is famously the thing nobody patches, and this is the answer to it.
The second engine does not start its update until the first has proven it came back — and proven means it restarted a service, ran a job, and reported out. Not that it answered a ping. If the first fails that check, the second stays where it is and rolls the first back.
And the second lags the first by a declared interval. A defect that surfaces after four days passes every health check ever written. The lag is what catches it, and it is a deliberate trade: briefly out of step, never both wrong.
Genie operates it. The cluster protects itself.
You should not be editing a topology file. Adding a node, declaring a tolerance, planning an upgrade window, working out which collector is saturated and rebalancing devices across the ones you have — all of that is a sentence you say and a change he shows you before he makes it.
Saying it is meant to be easier than finding it. The screens are all still there, and they are the fallback rather than the way in — because a screen nobody can find is worse than a sentence anybody can say, and an operator who cannot reach the model still has the screen.
He tells you when a poller is needed before it is needed, from the measured ceiling and your observed growth rather than a rule of thumb. He plans the upgrade — which night, what it interrupts, how long the lock should hold, and what to do if it does not take. And he explains the whole thing in plain English to whoever inherited it.
Tell him to fail a node over for maintenance and he does it — he plans it, says what it interrupts, and dispatches it through the same code path the automatic failover uses. What he is not is reachable by the automatic path. When the cluster loses its leader at three in the morning, quorum and fencing and promotion happen in code with no inference anywhere near them. A machine that has to ask a model what to do has already lost the property that makes it trustworthy.
The same line runs through every part of this. He cannot raise a tolerance to make a prediction go away. He cannot divide collection work by client. He cannot add a verb to the service engine, and he cannot pass an engine through its own health gate.
He proposes and a person commits. That is the same contract as every other thing he does here — and it is why the answer to what if it is wrong is never you have two primaries.
And a second engine is capacity as well
Collection is what decides how many assets a box carries, not inference. A practice at five thousand assets asks a handful of questions a day and polls five thousand devices every cycle. A second collector moves that ceiling, and we publish what we measured rather than what we projected.
Work is divided by device, never by client. No client, site or tenant ever lives on a node — a split by client would be four data models arriving by a different road, which is the thing this whole platform exists to avoid.
What this costs you
Nothing, and it is not a tier. It is off unless you turn it on, turning it on shows you the exact record that would be sent, and an air-gapped deployment runs the maintenance locally and sends nothing at all.
A vendor who can only help the customers who opted in has built something else. Support works the same either way.
Try it on your own estate
Thirty days, read-only, on your own hardware. No card, no call, and nothing to uninstall if you walk away.