Local changes to the ReefRhythm-SmartDoser firmware running the reef doser (see the Acropora page for the tank itself). This tree is a local working copy, not a git fork -- these fixes exist only as the comments/code in this checkout, found and fixed against real dosing behavior on this tank's own hardware (MKS SERVO42C stepper drivers, 750L reef system). Recorded here since there's no git history to carry them.
Firmware Changes
Dosing math (src/lib/stepper_doser_math.py)
- Low-flow calibration extrapolation could silently overdose. Below the lowest measured calibration point, the polynomial fit was extrapolating blind -- confirmed on real calibration data that it can swing back up instead of continuing toward zero. That defeats
move_with_rpm's low-flow runtime-shortening safety check entirely: a near-zero flow request could get an RPM at or above the achievable minimum instead of below it, so the dose ran at full requested duration instead of a shortened one -- a silent overdose. Fixed by anchoring a real (0 RPM, 0 flow) point -- a peristaltic pump at 0 RPM delivers exactly 0 flow, a physical fact the fit had no way to know on its own. - Extrapolated flow curve could decrease with increasing RPM. A degree>1 polynomial fit can wiggle (local extrema), producing a flow-rate curve that dips as RPM increases even though real flow never decreases with RPM. Forced non-decreasing via a running max.
- RPM-snap without runtime rescaling could silently overdose.
find_combinationsnaps a requested RPM to the nearest achievable one in the calibration table -- for a low-flow dose this is very often higher than requested (there's a minimum achievable RPM, ~0.593 on this table). Running that higher RPM for the original, unadjusted runtime overdoses, since delivered volume scales withrpm * runtime. Fixed by rescaling runtime down proportionally whenever the snap goes up. Deliberately one-directional: if a request's RPM is above the table's max (e.g. an unvalidated/runor MQTT command with a typo'drpm=99999), the snap goes the other way and scaling runtime bydesired/achievedwould be >1, amplifying runtime by the same factor instead. Failing safe there means delivering less than requested at the achievable top speed, never risking an unbounded overdose in the other direction. - Zero-step doses read as a driver fault.
calc_stepstruncates toint-- a request small enough that the runtime-shortening above collapses to well under one step's worth of rotation rounds down to exactly 0. Sending a real "run 0 steps" command to the driver isn't a no-op on this hardware: it came back as an outright failure (rejected reply or no completion confirmation within timeout), indistinguishable from a genuine driver fault, and triggered the same "DOSE FAILED -- driver unresponsive" alert for a request that was simply too small for this pump/duration to resolve. Fixed by checkingsteps <= 0and failing cleanly before any UART traffic at all.
Stepper driver protocol (src/lib/servo42c.py)
- CRC check used decoded values instead of raw bytes.
read_angle_errordeliberately does NOT usecheck_crc=True: the existing CRC check summed the already-decoded field values, not the raw bytes actually transmitted -- only correct for single-byte fields.erroris a signed 2-byte value (can be negative), so that comparison would not reconstruct the driver's real byte-level checksum and would reject valid replies. A correct fix needs the check computed overraw_databeforestruct.unpack, not after. - Delayed "straggler" reply corrupting the next command. Config changes (mode switches especially) can trigger a delayed side-effect reply/reset on the driver that arrives after the bounded read already gave up -- confirmed by observation: a dose sent immediately after a mode/current-set call failed, then an identical dose sent right after that succeeded cleanly. Fixed with a short settle + buffer flush before
send_set_command/release_protectionreturn. - Missing flush before make_steps when called with stop=False.
move_with_rpmcallsmake_steps(..., stop=False), which otherwise skips the only buffer-clearing step (stop()'s own flush) -- without an explicit flush here, a previous dose's "run complete" reply that arrived after its completion-wait had already given up would still be sitting in the buffer and get misread as, or corrupt, this command's own reply. - Fast doses could hide their own completion reply. Low-flow doses are very often RPM-snapped up (see #3), and for a big snap this can shrink actual motor runtime to a fraction of a second -- fast enough that the driver's "complete" reply can already be sitting in the buffer alongside the "starting" ack, in the very same
read_raw()call. Since that call already drained the buffer, waiting again afterward would never see it. Fixed by checking forcomplete_patternin the initial read first, before waiting for anything new. - Negative step count could raise an uncaught OverflowError.
steps.to_bytes(4, 'big')is unsigned -- a negativestepsvalue (reachable if an RPM-snap ratio or a malformed schedule entry ever produces one upstream) would raise uncaught instead of failing the dose cleanly. Fixed with an explicitsteps < 0guard.
OTA reliability (src/boot.py, scripts/init.sh)
- Version-comparison gate could leave files stale after an interrupted extraction.
version.txtis only ever written after a full extraction completes. If an extraction is ever interrupted partway (an external reset mid-boot, a crash), a later boot's version comparison could wrongly conclude nothing had changed and skip re-writing files that were never actually updated. Fixed by always re-extracting unconditionally on boot -- costs a few extra seconds every boot but guarantees static files can't get silently stuck out of date. - A silently-empty RELEASE_TAG made real OTA updates no-op.
RELEASE_TAGdrives boot.py's "did the app actually change" check and is shown on the OTA page / published over MQTT. If two builds in a row both had an empty tag, boot.py saw no version change and skipped re-extracting the frozen app -- so a real code fix could survive a "successful" OTA update without ever actually taking effect. Fixed ininit.shby defaultingRELEASE_TAGto a timestamped<version>-manual-<UTC timestamp>value whenever it isn't explicitly set by CI.
New features (not bug fixes, src/web.py, src/static/ota-upgrade.html, src/static/doser.html)
- GET /storage-status -- read-only endpoint exposing each pump's
container_size_ml/remaining_ml/dosed_since_refill_ml, reflecting the samestoragedict already used internally for the "container empty" alert. Added so cumulative dosing-over-time could be shown on this tank's own webpage without guessing it from the schedule alone -- a missed/failed dose or a manual off-schedule run wouldn't show up in a schedule-based estimate, but does here sinceremainingonly ever decrements on an actual completed dose. No side effects, no new write paths. - Cancel Download button on the OTA page. Previously there was no way to stop an in-progress OTA download short of power-cycling the device. Added
POST /ota-cancel, which sets a flag the existing download loop in/ota-upgradechecks once per chunk. Safe by construction, not just by care taken:OTA.write()only ever touches the inactive partition, and onlyOTA.close()callspart.set_boot()-- breaking out of the loop beforeclose()leaves the currently-running (old, known-good) firmware as the boot target no matter how much or how little was written. The one thing that did need explicit handling: dosing is paused (scheduled_jobs = []) the moment an OTA starts, and a cancelled download now restores it viaupdate_schedule(schedule), the same as the existing failure path -- without that, a cancelled OTA would have left real dosing silently disabled with no reboot to ever bring it back. - "Dosing now" indicator on doser.html. A pulsing indicator next to the currently-selected pump's name. Deliberately does NOT touch
stepper_run(the actual dosing execution path, already the source of several of the bug fixes above) -- two independent, purely client-side/frontend sources instead: scheduled doses are inferred by polling the existing, unmodifiedGET /schedule-diagendpoint every 5 seconds (already exposes each job's exact computeddose_timesand the device's ownseconds_since_midnight); manual doses are shown for the exact duration of the "Start Dosing" button's ownfetch('/dose')call, which already only resolves once the pump has actually finished (or failed) -- true start-to-finish coverage, not a guess based on the entered duration. Good enough for "does this look like something's happening right now", not a safety signal. - Settings save no longer reboots for cosmetic changes.
/settingspreviously called a hard device reset unconditionally on every save, purely as a byproduct of rewriting every config field in one shot -- renaming a pump bounced the whole live doser exactly the same as changing WiFi credentials. Only WiFi credentials, hostname (mDNS re-registration) and MQTT broker/login genuinely require the re-init this firmware only does at boot, and pump count changes fixed-size hardware arrays allocated at import time -- everything else (pump names, color, theme, notification settings, current limits, direction inversion) is now rebound live in place instead of waiting for a reboot. - "Dosing now" indicator could miss short doses. The scheduled-dose polling above ran every 5 seconds -- fine for longer doses, but this tank's shortest scheduled dose (Magnesium, ~15-21s) could have up to a quarter of its actual run window missed at either edge. Tightened to a 1-second poll interval; manual doses were already unaffected.
- Scheduled dose-amount ramps. A schedule entry can now carry a start date, end date, start amount and end amount; the doser linearly interpolates the actual dosed amount between them by calendar date, holding at the start amount before the start date and permanently at the end amount after the end date. Symmetric by construction -- an end amount lower than the start amount ramps down, not just up. A new background check runs once a minute, recomputing the live schedule whenever the calendar date rolls over on a ramping job (a no-op for every other schedule entry). Added because gradual dosing changes -- this tank's own carbon-dosing ramp-up, ongoing Ca/Mg/Alk corrections -- previously required manually recalculating and re-pushing a new amount every few days by hand. The schedule editor on doser.html also gained a "Ramp dose amount over time" option (start date, end date, end amount) when adding a new dose, which scales duration together with amount to keep the pump's flow rate roughly constant across the ramp, the same approach used for every manual dose correction on this tank so far.
- Dose-log write failures could go completely unnoticed. Every dose gets persisted to an on-device history file for later review, but a write failure there was only ever visible on a serial console -- confirmed on this tank when real dosing kept running correctly (notifications went out the whole time) while the history file itself had a multi-hour gap with zero trace anything had gone wrong. Fixed by routing a write failure through the same notification channels already used for dose-failure alerts, rate-limited to once per failure episode rather than once per dose (this runs on every single dose across every pump, many times a day), with a matching one-time "recovered" message once writes succeed again.
- New
GET /fs-status. Read-only endpoint exposing the doser's internal filesystem usage (total/used/free bytes), added to check whether a near-full filesystem could explain the write failures above. On this tank it was ruled out (47.9% used at the time). - Ramping schedule points now show today's actual dose in the info modal. The dose-ramp feature above means the amount a pump actually runs today can differ from what's shown in the raw schedule list -- there was no way to see today's real interpolated dose without querying the device's diagnostics directly. Clicking a ramping point's info popup now shows "Today's actual dose" alongside the existing ramp-range text, fetched once at click time rather than continuously polled -- no extra network traffic for ordinary fixed-amount schedule points.
Dose diagnostics and an advanced settings page (src/web.py, src/lib/servo42c.py, src/static/advanced-settings.html)
- Dose failures were logged without saying why. The driver code already distinguished several real failure causes internally -- no reply at all, an explicit rejection, a run that started but never confirmed completion, the low-flow implausible-RPM safety refusal, a negative step count -- but only ever printed which one to a console nobody has open. The persisted dose-log history collapsed all of them into a bare
FAILED, with no way to tell afterward which had actually happened. Fixed by threading a reason string through every failure path into the existing dose-log line, plus a fallback note for any future un-annotated failure path, so a gap in this coverage shows up as "reason not captured" instead of silently reproducing a blankFAILEDagain. - Per-dose waits could stall the web server.
set_mstep's retry loop andmake_steps's initial "run starting" acknowledgement both used a blocking,time.sleep_ms-based wait rather than the event-loop-yielding one already used for the (much longer) completion wait -- meaning the web server, WiFi, and every other pump's schedule could stall for up to roughly 900ms in the worst case, on every single dose, not just an occasional one. Converted both to the same async, yielding wait pattern already proven on the completion wait. - New /advanced-settings page. The MKS SERVO42C protocol documents a considerably larger command set than this firmware ever exposed -- working current, microstepping, position-loop PID gains (Kp/Ki/Kd), acceleration, max torque, locked-rotor protection, on-demand calibration, and more, all reachable over the same serial bus already used for dosing, with no UI to reach any of it. Added a dedicated page exposing these settings grouped by real consequence (cosmetic, operational, dangerous) rather than a flat list, each with a plain-language explanation of what it does and what breaks if it's set wrong. Dangerous settings require typing the target pump's name before the send control activates, and the same confirmation is enforced server-side too, independently of the page's own JS -- the request itself is refused without it, not just hidden behind a UI gate.
- Work Mode, Serial Baud Rate, and Restore Factory Defaults deliberately left out. All three can take a driver out of CR_UART/38400 entirely, and the only way back is the driver's own physical OLED menu -- not a viable recovery path on this installation. For every other dangerous setting on the page (bus address, acceleration) a real software recovery exists; for these three there isn't one, so a wrong value here wouldn't just be risky, it would be unrecoverable. Left off the page entirely rather than hidden behind a warning.
- Raw bus-address recovery. If two drivers ever end up answering the same address -- for instance if one was ever reconfigured via its own physical menu to an address another pump already uses -- the normal pump-name selector becomes actively misleading: it addresses a pump by its intended identity, not by whatever address its driver is actually sitting at. Added a direct raw-address override that talks to a specific bus address regardless of which pump is supposed to live there, so recovering from a collision doesn't depend on guessing which pump name happens to reach the colliding driver.
UART bus-sharing race condition, introduced by the wait fix above (src/web.py, src/lib/servo42c.py, src/lib/stepper_doser_math.py)
- Making dose waits non-blocking reintroduced a race the old blocking version had accidentally prevented. Item #22 above fixed the web server stalling during a dose by making
set_mstepandmake_steps's initial acknowledgement wait yield the event loop instead of blocking it. What wasn't obvious at the time: the old blocking wait had an unintended side effect, freezing the single-threaded event loop for a dose's entire duration also happened to block every other UART-touching route (/pump-diag,/pump-config, the advanced-settings page) from ever running concurrently with a dose -- and all nine stepper drivers share one physical UART bus. Removing that freeze let diagnostics and config calls interleave their bytes with an in-progress dose's own traffic, corrupting both sides of whichever two commands collided. Confirmed by measurement on this tank: the dose failure rate roughly quadrupled, from 7.9% (20 of 252) in the days before this change to 34.5% (20 of 58) after, hitting nearly every pump. Fixed permanently -- not by reverting the event-loop fix, which stays in place -- by routing every one of these previously-direct driver calls through the samecommand_bufferqueue/dose,/stop, and/runalready used to serialize access to the shared bus, so diagnostics and config changes now queue behind an in-progress dose instead of racing it.
Reading replies on a shared half-duplex bus (src/lib/servo42c.py)
- The driver's own transmission was being read back as the driver's reply. All nine drivers share one single-wire half-duplex UART, so every byte the ESP32 writes also lands in its own RX buffer. Nothing in the read path distinguished that echo from a genuine reply, so a command whose driver never answered still "succeeded" -- it just parsed its own outgoing bytes. The arithmetic made it unambiguous:
read_protection_statereturned 62, which is exactly the0x3Efunction code it had just sent, andread_angle_errorreturned the bus address plus 14393, where 14393 is0x3839-- the0x39function code followed by its checksum byte. Fixed by recording each command as it is written and stripping that exact prefix from whatever comes back before parsing. This also re-labelled a class of failures that had previously been reported as bogus sensor values rather than as "no reply". - A failed read was unpacked into three variables, returning 500 instead of "no reply".
read()signals failure by returningFalse, butread_angle_error,read_encoderandread_pulsesimmediately unpacked its result into multiple names. With a driver physically disconnected -- or simply not answering -- that raised aTypeErrorinside the request handler, so/pump-diagserved an opaque "Internal server error" exactly when it was needed most. Each now checks forFalseexplicitly and reports no-reply as a value.
Config and diagnostics could freeze the whole device (src/web.py, src/lib/servo42c.py)
- Every driver config and diagnostics call blocked the single-threaded event loop for its full wait. Fixing the echo misread above had a side effect: a call that used to return instantly on its own echo now correctly waited out its entire bounded timeout. Because these methods still used
time.sleep_msand ran directly inside the event loop, one unresponsive driver could stall the web server, the scheduler and every other pump for up to ~1.5 s per call. Added*_asynccounterparts that yield the loop during their waits -- the same pattern the dosing path already used -- and routed every diagnostics and config route through them. The original synchronous methods are kept deliberately, becauseload_configs.pycallsset_current()at boot before an event loop exists. - Those new async calls were silent no-ops for two builds, because MicroPython coroutines have no
__await__. The dispatch helper decided whether to await a result withhasattr(result, "__await__"), which is correct on CPython. In MicroPython a coroutine is a generator and has no such attribute, so the test was always false, the coroutine was never awaited, and every configuration write silently did nothing while reporting success./pump-diagmade it visible by returning the string<generator object 'read_angle_error_async'>as a pump's angle error. Fixed by testing against the real generator type as well as__await__. Worth recording as a caution: a CPython idiom that merely returns the wrong answer here, rather than raising, produced two builds in which an entire feature appeared to work and did not.
Dose acknowledgement window and OTA download speed (src/lib/servo42c.py, src/web.py)
- The run command's acknowledgement had the tightest timeout on the bus and no retry behind it.
make_stepswaited 150 ms for the driver to acknowledge a run command and abandoned the whole dose if it did not arrive, while a neighbouring exchange (set_mstep) retried the same kind of read five times. Widened to 400 ms, which costs nothing when the acknowledgement is on time and carries none of the double-dose risk a retry on a motion command would. See item #34 for what was actually making those acknowledgements late. - The OTA download spent about 75 seconds doing nothing. The firmware download loop yielded with
await asyncio.sleep(0.1)once per 4096-byte chunk, so a ~3 MB image paid a fixed 100 ms penalty roughly 750 times over. Changed toawait asyncio.sleep(0), which yields to the event loop just as effectively without the delay. Chunk size is deliberately unchanged, so the yield frequency -- and therefore responsiveness during an upgrade -- is identical.
The scheduler silently skipped doses whenever the event loop ran late (src/web.py)
- Matching the current second exactly meant a dose was lost outright if no tick landed on its second. Both scheduler loops ran
if now in job["dose_times"]once perawait asyncio.sleep(1). But that sleep guarantees at least a second, never exactly one, so whenever the loop had other ready work a tick advanced well over a second of wall clock and stepped straight over the seconds in between. A dose due on a skipped second was not retried, not logged, and raised nothing -- it simply never happened. Becausescheduler_tick_countincrements once per tick, dividing it by the device's own clock measures this directly: 0.705-0.746 ticks/s with two browser pages open (25-30% of seconds never examined) against 0.983-0.994 with none. Matching the schedule against the dose log over one 11.6 h window found 89 doses scheduled, 68 logged and every log row accounted for -- so 21 doses, 24%, about 106 mL, had never run, with the log sitting at 68 lines against a 200-line trim limit so nothing had been discarded. The decisive detail: two pumps that share a dose second disappeared together on all five occasions, which no per-pump comms or driver fault can produce. Fixed by firing on the half-open interval since the previous tick instead of on an instant, which makes dispatch correct however far behind the loop runs. The interval is committed before any dose is dispatched, so a dispatch that raises cannot reopen it and dose twice; it is exclusive at the previous second and inclusive at the current one; it handles the midnight wrap; the first tick after boot starts one second back rather than replaying the day; and intervals wider than 180 s are discarded rather than honoured, since an NTP correction backwards reads as a ~86400 s jump forward and replaying a long stall would dump a burst of doses at once. Simulated against this doser's own 192-dose schedule, the new logic delivers 192/192 with no duplicates at 0, 20, 430, 1500 and 5000 ms of lag per tick, where the old exact match delivers 134/192 at 430 ms. - A once-a-second poll for a cosmetic indicator was what made the loop run late -- and it was also failing doses. The "dosing now" dot from items #14 and #16 polled
/schedule-diagevery 1000 ms from each open page and compared times client-side. That endpoint sorts every job'sdose_times(192 entries across 7 jobs on this doser) and serialises ~2.6 KB on every call, so each open page bought the device that work once a second. The interval fix above stopped this costing whole doses, but it did not stop it failing them on the run-command acknowledgement of item #31: over one 5.5 h run, 1 of 16 doses failed with pages open and 26% of seconds skipped, against 0 of 33 with pages closed and 1% skipped. In other words, watching the pumps degraded the pumps. Replaced with a newGET /pump-running?id=Nthat returns a single boolean -- 18 bytes against 2629, with no sort and no large serialise -- and the page polls that at the same 1 s cadence. Semantics are deliberately unchanged: still inferred from the schedule, still unable to tell a real dose from one that silently failed to start, still blind to a manual off-schedule dose. Computing it server-side also fixed a real edge case the client-side comparison got wrong, a dose that starts before midnight and is still running after it. Measured afterwards: a simulated page polling once a second costs no measurable lag at all (+8 ms/tick, identical to an idle device), and adding a live/dose-ssestream on top brings it only to +17 ms, against +403 ms before. That last measurement is worth stating plainly, because it was deliberately a one-change-at-a-time test: the per-second poll was essentially the entire cause, and the SSE stream -- the original suspect -- contributes almost nothing.
Current build hosted at /downloads/firmware_smartdoser/ is built from this exact fixed source tree (verified by sha256 match against the local build output before publishing).