Your Driver List Is Not a Lever List
Prediction, intervention, and what-if are three different questions, and only two of them need to know how your business is wired. Here is how to tell which one your report answered.

The list of drivers sitting under your forecast is a different object from a list of levers. A system can match the strongest predictive baseline on what happens next quarter and still be wrong about what to change, because those are two different questions, and only one of them requires knowing how the business is actually wired. The instinct that tells an operator which lever to pull is usually the sound part of the exercise. What has been harder to check is whether the analysis on the page answered the question that was asked or a neighboring question that looks identical in a chart.
The Two Questions on One Screen
Say the managed services line gave up two points of gross margin this quarter and the report attributes the movement to overtime concentrated on three accounts. Two questions follow immediately, and they arrive in the same conversation with the same air of authority. The first is what the number does next quarter if nothing changes. The second is what happens to the number if the staffing model on those three accounts changes. A third arrives later, usually in a board review or a diligence session: what this quarter would have looked like if the repricing had happened in January.
Those questions ask for different things. The forecast needs relationships that hold well enough to extrapolate, and it is indifferent to which way the arrows point between overtime, escalation volume, and contract structure. The change question is not indifferent, because a factor that travels alongside margin and a factor that produces it call for entirely different actions. The what-if question asks for more still, holding everything that already happened fixed while one decision is rerun. Most operator-facing analysis answers the first question competently and then presents the result in language that sounds like it settled the second.
The Three Levels Come Apart
Meaningful evaluation makes the split checkable, because it runs the question types against a system whose wiring was known in advance.
One recent experiment supplied a tabular model with the true causal structure of a system and then measured all three question types against the same underlying data, which turns a distinction that usually stays theoretical into something you can read off a table. The setup was synthetic and deliberately unfamiliar to the model: the comparison ran on out-of-distribution structural causal models with 1,024 fit samples, scored by mean squared error, where lower is better. Handing the model the true structure moved interventional error from 0.765 to 0.739 and counterfactual error from 0.607 to 0.572, while observational error sat at 0.537 without the structure and 0.538 with it. The mechanism behind that split is plain: causal direction does not affect the observational distribution, so knowing the wiring buys nothing on a question that only asks what tends to move together.
Read that as procurement guidance rather than as method and it gets uncomfortable. A system can sit at parity with the best predictive baseline on the forecast and carry no warrant at all for the recommendation printed underneath it. Telling a system what you already know about how your business is wired will do nothing for next quarter's projection and quite a lot for the decision you make about it. The tests were synthetic, so the transfer to a services P&L is an argument rather than a measurement, and the argument holds because the arithmetic of the three questions does not change when the variables become client counts and billable hours.
What an Honest Answer Says About Itself
The more useful part of that evaluation is what it did with the variables nobody measured. The structures were represented as acyclic directed mixed graphs, in which a bidirected edge stands for unobserved confounding, a common cause sitting outside the recorded data. Those confounding edges were recovered at 0.828 area under the receiver operating characteristic curve on the out-of-distribution structures at 1,024 fit samples. The figure matters less than the posture it makes possible. A system built this way can hand back an answer that states its own limits, including the case where the structure implied by your data cannot be pinned down and the honest output is the ambiguity rather than a confident factor list.
That is the register we hold QortexOS to, and it shows up as mechanism instead of a value statement. Cash forecasts arrive as a low, expected, and high range with covenant breach risk and what-if stress tests behind them. Risk predictions carry the specific factors moving them and a recommended action. Intelligence that is not confident enough is routed to a person rather than accepted, and the constrained optimizer names the limit that is holding profit back so the recommendation arrives with its boundary attached.
We do not ship causal structure discovery today and are not claiming it here.
The argument is about a question worth asking of any recommendation, including one of ours: which of the three questions did this answer, and what had to be true about the wiring for the answer to hold.
Thin History Flatters What Already Happened
There is a second result in that evaluation with a direct read for anyone running a young client relationship or a freshly acquired book of business. At small sample sizes the counterfactual prediction was biased toward the outcome that had already been observed, and the bias faded as the number of samples grew and the model learned the underlying system better. That is a property of sample size rather than a defect in the approach, and it translates cleanly. A what-if run against two quarters of history for a client onboarded in the spring will tend to agree with whatever already happened to that client. Run the same question against four years of history on a long-tenured account and the answer has enough room to disagree with the record.
For an MSP in the middle of an acquisition, that gap decides whether a diligence model reflects the acquired firm's operating reality or mostly restates its recent past in more confident language. The practical response is simple. Weight a recommendation by the depth of history standing behind it, and put that depth on the face of the answer rather than in a methodology page nobody opens.
The Check Before the Funding
Here is the short form of what to ask before a recommendation moves money.
Before you act on the answer
The operator move here is small and it holds regardless of which system produced the output. When a recommendation lands, ask which of the three questions it actually answered, and require the structural assumptions it depends on to be stated where the answer is, not buried a layer down. A forecast that is right about the number and wrong about the lever will still get funded and defended in a board meeting, and the cost surfaces two quarters later when the change produced nothing anyone can point to. Visibility of the numbers and visibility of the structure are separate deliverables. The second one decides whether the move you make is the move that mattered.
More from Insights

Forecast Error Follows the Operating Picture
Manufacturing forecasts at 10.2 percent error, services at 14.5. The four-point spread traces to seams between the systems a growing firm runs, and it gets paid for in margin.

Work Design Shows Up in Behavior First
The quiet over-performer is an early reading on the work design, and absenteeism and attrition are the expensive versions of the same information, arriving after the business has paid for it.

The Attack Surface Nobody Inventoried
A file-transfer service you never bought ships enabled on every phone and laptop in the building, answers unpaired devices in wireless range, and belongs, on paper, to nobody.
See sooner. Decide faster. Act with confidence.
Ask Which Question Your Forecast Answered
QortexOS the operating system for the modern MSP.