Notes from tracking down why both of this tank's 2x ATI Straton Flex 153 fixtures (150cm, driven by a custom Python solar-curve scheduler rather than the manufacturer app's own timeline editor -- see the Acropora page) stopped responding after several days of otherwise normal operation. Published because the root cause turned out to be inside ATI's own firmware, not this tank's automation script, and the failure mode (fixture's web interface silently dies, only a power cycle brings it back) is the kind of thing another owner would otherwise spend a while chasing as "my own script must be doing something wrong." These fixtures are out of warranty, which is the only reason SSH access into the underlying OS was used for this -- opening the unit and running a manufacturer diagnostic tool would obviously be the first step to try instead if warranty coverage still matters to you.
More generally: I paid for this hardware, and once something is out of warranty I think it's mine to open, inspect and fix as I see fit -- that's the whole point of actually owning a thing rather than renting it. If any of this is useful to you and you decide to try it on your own fixtures, that's entirely your call to make and entirely your own risk to carry: you are fully responsible for whatever happens to your own hardware, network, data or warranty status as a result. Not me.
Both fixtures' web interfaces stopped responding within about an hour of each other (2026-09-25, roughly 22:00 and 23:00 local), after both had been running continuously since a factory reset five days earlier. Each fixture kept answering ping the whole time -- the underlying Linux OS and network stack were fine -- but every TCP connection to ports 80 (web UI), 22 (SSH) and the fixture-to-fixture sync ports was refused. Only a power cycle brought either one back. The automated schedule push (a script run every 15 minutes) logged this as a plain connection-refused error and moved on, per-fixture, without affecting the other unit or its own next scheduled attempt.
ATI's own app has a "Settings → Download Support File" button, producing a .tar.gz with the device's logs, running process list, memory info and restart history -- the single most useful thing for chasing this, and worth grabbing from both fixtures the first time this happens on any similarly-affected unit, since the in-RAM log it's built from is wiped by the very power cycle that fixes the symptom.
The hardware underneath a Straton Flex turns out to be an Onion Omega2+ board (MediaTek MT7688, 124808kB RAM, no swap) running OpenWrt 18.06, with the actual fixture app being node --expose_gc server.js, started by procd and restarted automatically on crash -- but only up to a point (procd's own respawn setting for this service is respawn_threshold=15 respawn_timeout=1 respawn_retry=1: it gives up after repeated quick failures, logged in logread as "Instance ati2::instance1 is in a crash loop"). Support-file logs from a fresh boot showed the node process starting at roughly 102-118MB of virtual memory already, on a device with under 125MB total and no swap to fall back on.
The app's own config.js (extracted and read directly, not reverse-engineered from the minified client bundle) has a built-in memory watchdog: memwatchMaxMemory: 130, checked every memwatchInterval (5) minutes against the process's real virtual size via lib/memwatch/index.js, calling process.exit(0) the moment it's exceeded. With a ~102-118MB baseline against a 130MB limit, that left only 12-28MB of headroom for any long-run growth (session state, accumulated color-blend data from repeated schedule pushes, or a genuine leak -- the support file's single-boot snapshot can't distinguish which) before the watchdog fires -- and once it does, procd's crash-loop protection means nothing brings the app back on its own.
- Raised the memory watchdog's limit, 130MB → 150MB. A one-line edit to
memwatchMaxMemory in /www/ati/config.js on both fixtures (each backed up first). Note for anyone trying this from the app's own Settings page instead: the UI's save handler for this value (requests.config.js) is a no-op that always reports success without writing anything -- the only way to actually change it is editing the file directly and restarting the app (/etc/init.d/ati2 stop; ...; /etc/init.d/ati2 start, the exact sequence the fixture's own weekly cron job already uses).
- Reused the login session instead of logging in on every push. The scheduler script previously logged in fresh on every 15-minute run -- roughly 190 new sessions a day per fixture, never explicitly logged out (the fixture's web app exposes no logout endpoint to call). Each login is a small memory allocation on a device with very little headroom; the script now saves the session cookie and only re-authenticates when the fixture actually rejects it.
- Stopped rewriting the schedule when nothing changed. The previous behavior rewrote the fixture's entire stored timeline and ~97 custom colors (a ~155KB payload) every single run, whether or not the computed schedule had actually changed since the last push. The script now compares the newly computed schedule against what the fixture already holds (time, intensity within 0.1%, each channel within 1 unit) and skips the write entirely when they match -- in practice this now skips the vast majority of the 96 runs/day per fixture.
- Added an automatic recovery path. A separate check now polls each fixture's web port every 5 minutes; if it's been unreachable for three checks in a row (15 minutes) and SSH still answers, it restarts the app remotely using that same
/etc/init.d/ati2 sequence -- rate-limited to once per 6 hours per fixture, with a notification either way (recovered, or SSH is also unreachable and a physical power cycle is needed). Turns most future occurrences of this exact failure into an unattended ~30-second blip instead of a dark fixture until someone notices and walks over to unplug it.
- Staggered the fixtures' own built-in weekly restart. Both units ship with an identical
0 3 * * 0 root cron entry that restarts the app every Sunday at 03:00 -- harmless on its own, but it was landing on the exact same minute as this tank's own 15-minute schedule-push cron, meaning a scheduled push could occasionally race a self-restart. Moved to 03:22 and 03:37 respectively, clear of every quarter-hour boundary.
- Switched from a password to SSH keys for the automation's own access. The recovery check above and the memory-monitoring probe both need to log in to the fixtures directly; that ran on a stored password at first, now on a dedicated key instead (password login is left enabled on the fixtures themselves as a fallback, in case a factory reset or similar ever wipes the key). One snag worth mentioning: these fixtures run Dropbear v2017.75, which is RSA-only -- an ed25519 key is silently rejected outright, no error beyond "permission denied", so RSA is the one to generate if you're doing this yourself.
- Hardened the scheduler against its own upstream dependency failing. Unrelated to the crash above, but found and fixed the same week: the script's live cloud-cover forecast (used to dim the lights under real clouds) had a single point of failure -- if the weather API returned an error, the whole day's remaining schedule silently fell back to full clear-sky, then flipped back once the API recovered, producing a visible jump in the middle of the day for no reef-side reason. Now falls through: live forecast → last good forecast (if under 12h old) → a typical-day cloud pattern learned from that fixture's own logged history → flat clear-sky only as a last resort with no other data at all. Also now averages the forecast across two separate reference reef locations rather than one, so a single location's own bad data doesn't skew the result either.
- Replaced the login-form/cookie session with a real token-authenticated API. The session-reuse fix above helped, but the real problem was structural: the fixture only allows one active login session per user, so a manual browser login (checking on it, or the ATI app itself) could silently knock the scheduler's own session out with a 401, and vice versa -- exactly the kind of hard-to-reproduce failure that looks like "my script must be buggy" until you happen to catch it live. Patched
Requests.prototype.auth in /www/ati/requests.js -- the one function every authenticated route already runs through -- to accept a per-fixture bearer token (stored alongside memwatchMaxMemory in config.js) ahead of the existing session check. Present and correct, it authenticates the request immediately with no session created and no interaction with the single-session state at all; absent, everything falls through to the original logic completely unchanged, so the web UI and the official app's own password login are untouched by this. The scheduler now sends Authorization: Bearer <token> on every request and holds no cookie at all -- there's no session left to be knocked out of.
- Started keeping the fixtures' own log output off-device, in real time. The support-file download and the recovery/monitoring scripts above all ultimately depend on
logread's in-RAM ring buffer -- which is exactly what a power cycle wipes, the moment the very symptom it would explain (a watchdog trip, a crash loop) has just happened. Pointed each fixture at a small syslog receiver on the same server as this website, using OpenWrt's own built-in remote-logging config (log_ip/log_port, no extra software needed on the fixture side) -- so those lines now land in a dated file per fixture that survives any reboot, with a daily job clearing anything older than 30 days.
Strong recommendation: update to the latest firmware available for your fixture, including a beta release if that's what's currently offered. Ordinary software hygiene: bug fixes and improvements only reach a unit that's actually been updated, and checking for and installing the newest available version -- rather than assuming "it's been working fine" is a reason not to -- is one of the simplest things an owner can do here. Firmware downloads: ATI Aquaristik's Straton Pro/Flex support portal (the manufacturer's own site, EU/German), or ATI North America's support page under the Straton LED category.
The single biggest improvement available for a fixture like this, beyond anything above, isn't a firmware patch at all: put it on its own isolated (V)LAN, reachable only from one designated administration machine (a single PC, laptop, whatever you actually use to manage it), not the general home network. Everything found while chasing this crash was reachable to anything already on the same LAN segment as the fixture -- other IoT devices, guest devices, anything that's ever had the Wi-Fi password. None of the fixes above change that exposure; they make the fixture more reliable and this tank's own automation more robust, but the fixture itself is still only as trustworthy as everything else sharing its network. Segmenting it away from the rest of the home network, with a narrow, explicit allow-list for only what actually needs to reach it, lowers the real-world risk of everything in this writeup (and the vulnerabilities we're not detailing here) far more than any single code change could.
While going through the app's own code to diagnose the crash above, a couple of other things turned up that aren't covered in this writeup. ATI: if that's you, please get in touch via my contact page and I'll share the details directly.