Sensors Still Need a CPU.
IoT on a remote oil field still has to run on nearby computers. The paper is about which CPU should run each step of a fire-detection workflow before the deadline, written down so a solver can prove the answer.
Sensors on an offshore oil platform do not care that a cloud region exists. They produce smoke readings, heat readings, camera frames. If a fire starts, a chain of work has to run now: clean the signal, decide if it is a fire, figure out where, simulate how it spreads, push the alerts. Each step needs the one before it.
The internet of things is a polite name for that pile of sensors. The impolite part is computation. Those streams still have to run on CPUs, the CPUs nearby are small, and the cloud is often a satellite hop away, or gone. I spent a summer writing that placement problem down as a mixed-integer linear program, for federated fog systems, and we published it at the IEEE Cloud Summit 2025.
This piece is the motivation. The other one is what happened when the exact answer stopped finishing.
Why the cloud is the wrong computer
Industry 4.0 is a conference name for something simpler. Factories, rigs, and remote platforms now run on autonomous, data-driven workflows. The paper’s examples are offshore oil fields, space stations, and underwater robots. They share three facts: connectivity is intermittent, people on site are scarce, and a lot of the work is safety-critical. Disaster response, predictive maintenance, real-time environmental monitoring. You cannot wait for a round trip to a data center if the link is down and the fire is not.
Fog computing is the unsexy answer. Put small computers, mini data centers, next to the sensors. Process there. Do not ship every frame to the cloud.
One fog box is not elastic. During an oil spill or an equipment failure you may have several of these workflows at once, all on a deadline. A single node’s CPU fills up. Nearby fog systems can federate, share compute, and look a little more like a cloud. Sharing also creates the actual problem: heterogeneous machines, different speeds, and a communication delay every time you send a task’s output to a neighbor.
Resource allocation, in this paper, is not a metaphor. It is: which fog node’s CPU runs which task, so the whole chain still lands before the deadline.
The workflow is a graph, not a to-do list
Fire detection in the paper is a Directed Acyclic Graph. A DAG. Each node is a task. Sensor data collection, smoke and heat analysis, fire localization, alert dissemination, plus later steps like evacuation planning and fire-expansion simulation. An arrow means: this step cannot start until that one has finished and handed over its data.
A to-do list you can shuffle. A DAG you cannot.
Put localization on a fast CPU that is far from the node that just finished detection, and you pay a wait for the data to arrive. Put everything on one box and that CPU is the bottleneck. The schedule is a trade the whole way down.
We modeled a synthetic DAG on that real fire-detection workflow and used it to test placement. The industrial constraint underneath is deterministic execution: Industry 4.0 workloads want the chain to finish on time, not “soon, usually.”
A mixed-integer linear program, in English
A linear program is a very strict packing problem. You have knobs (numbers), a score you want to maximize, and rules that stay straight. Double a task’s work, double the time it takes. No hidden curves.
Integer means some knobs are not allowed to be 2.7. They are 0 or 1. Yes or no. Mixed means you have both kinds.
In this model the yes-or-no knobs are placement. Does task i run on fog node a. Each task runs on exactly one node. You cannot start it on one CPU, pause it, and finish it on another. The fraction knobs are how much of the task to actually run, between 0 and 1. Run it fully and it scores its full accuracy. Run it partway and accuracy drops linearly. The score we maximized was average accuracy across tasks, as long as every task still finished by the deadline.
The rules are just physics written so a solver can read them. A task cannot start until every task feeding it has finished, plus the communication delay if those two CPUs are not the same machine. Nothing crosses the deadline. If a node is not allowed to take a task, it does not.
When the solver returns, the answer is not a vibe. It is the best assignment that exists under those rules, and you can prove it.
The paper also had to turn the first version of the math into a true MILP. Some of the natural rules multiply variables together, which a linear solver will not take. We rewrote those products into equivalent linear constraints. Same solutions, readable by the solver. The other post has the Pyomo.
Approximate on purpose
People hear “partially execute a task” and think we shipped wrong answers for fun. In a fire on a rig, a coarser spread estimate you get in time beats a perfect one that lands after the decision is already made. The paper is explicit that how much accuracy you can give up depends on the job. Early anomaly detection might accept on the order of a 10 to 15 percent drop if the response is much faster. Mission planning or control may need almost all of the fidelity. Quantifying those thresholds is still open. The model puts the trade on the table instead of pretending full accuracy is free.
Standalone fog is inelastic. Federation is how you buy elasticity without the cloud. Approximate computing is how you buy time when even the federation is tight.
Why this was worth a paper
Scheduling dependent tasks across machines like this is NP-hard. Ten tasks solve instantly. A hundred make the exact solver sweat. A thousand, which is the size a real system wanted, and the optimum is technically on its way and practically never coming. We paired the MILP with a greedy scheduler that assigns each ready task to the CPU that can finish it soonest, then shrinks durations together if the whole graph still misses the deadline. Greedy is worse on accuracy. It finishes.
We swept deadline slack, number of fog nodes, and graph size. More slack helped both; the MILP stayed ahead; greedy flattened once deadlines got loose. More nodes helped the MILP every time; greedy stopped improving past eight. On small DAGs the exact model is simply better. On large ones it falls off a cliff because it stops returning. Greedy holds a lower line out to a thousand tasks.
To our knowledge this was the first systematic comparison of MILP and greedy task allocation for DAG scheduling in federated fog. The future work in the paper is the honest part: run it on real offshore and industrial IoT deployments, not only synthetic graphs, and look at relaxations and metaheuristics so the exact model can get closer to the sizes that matter.
I flew to Madrid for this, NSF IRES, with Gayani Gupta as first author, and Mohsen Amini Salehi, Antonio Fernández Anta, Sanjukta Bhowmick, José Aguilar, and Mathieu Kourouma. IMDEA, the University of North Texas, Berkeley, Southern University. The trip ran on NSF IRES grants. I owned the formulation in Pyomo. The motivation was never “publish a MILP.” It was: the sensors are already out there, the CPUs next to them are small, and if the cloud is gone you still have to decide who runs the fire.
FAQ
- Why not just send the sensor data to the cloud?
- Because the paper’s setting is the case where you cannot. Intermittent or missing connectivity, harsh remote sites, strict deadlines. Fog exists so the workflow can finish locally.
- Is a mixed-integer linear program just a fancy spreadsheet?
- Close. A spreadsheet you nudge. A MILP you state the decisions, the score, and the straight-line rules, then a solver searches the legal assignments and returns the best one it can prove. The integer part is the yes-or-no placement. The mixed part is also allowing a fraction of each task to run.
- If greedy scales, why write the MILP at all?
- Without it you have no ground truth for how much accuracy the fast method is leaving on the table. On small and mid-size graphs the exact model is better. Plenty of real workflows live there. The companion post is the scaling story.
Amisha
Filed under Research · Distributed Systems