LoadingSoftware that carries weight.ONYXERA TECHLoading route
SLA-backed support for the systems we built — and the ones you inherited from somebody else. Monitoring, security patching, incident response and a standing engineering retainer, priced per month.
Handover is a formality. The team that shipped it stays on the account, so there is no rediscovery phase and no context tax.
An agency wound down, a contractor moved on, or a founding engineer left. We audit, document and take the pager.
Half on the legacy stack, half on the new one. We keep both alive and instrumented until the cutover is actually finished.
Fixed monthly pricing, a written response time, and an engineer who already knows your codebase. Move between tiers with 30 days notice — no rebuild, no re-onboarding.
Keep it healthy.
Monitoring, patching and defect fixes for a product that is live and behaving. The floor under your uptime — not a growth plan.
Keep it moving.
Everything in Essential, plus a standing engineering retainer so the roadmap keeps moving in the months between build projects.
Keep it accountable.
A dedicated pod, a contractual SLA and 24/7 paging for systems where an hour of downtime already has a dollar figure attached to it.
Running something that does not fit a tier — a regulated workload, a fleet of white-label deployments, a migration that needs two teams? We write custom plans against the same SLA language. Send the details to support@onyxeratech.com and you will get a scoped answer, not a brochure.
Repository and infrastructure audit, dependency and CVE inventory, a written risk register ranked by blast radius.
Synthetic checks on your critical paths from four regions, structured logs, p95 and error-budget dashboards, alerts routed to a named rotation.
One runbook per failure mode, short-lived credentials from your vault, escalation tree agreed with your engineering lead.
Severity definitions signed, response times live, first patch window scheduled. From here the SLA is contractual, not aspirational.
Scripted checks walk your critical paths every 60 seconds from four regions. We find out before your users do, which is the entire point.
A weekly window for dependencies, base images and runtimes. Anything scored CVSS 7.0 or higher is remediated inside seven days.
A named commander, a live status thread in your channel, and a blameless written postmortem within 48 hours of resolution.
p95 latency and Core Web Vitals are tracked per release and enforced in CI, so a regression fails the build instead of shipping.
Encrypted backups with point-in-time recovery, plus a scheduled restore rehearsal so the runbook is proven rather than assumed.
One page: uptime, error budget burn, open defects, hours used, and the three things we would fix next. No dashboard you have to remember to open.
Every plan ships with the same severity definitions. What changes between tiers is how fast a human picks it up — and whether that human is awake at 3am.
Production is down or unusable for all users, or customer data is at risk. No workaround exists.
A core workflow is broken or badly degraded for a large group of users. A workaround exists, but it costs them.
A secondary function is broken, or a defect affects a small subset of users with a straightforward workaround.
Cosmetic issues, copy corrections, questions and improvement requests with no measurable user impact.
The console below is the same view our rotation works from: rolling 30-day availability, live p95, queue depth and the alert stream. Every incident on it has a written postmortem attached.
Representative view of the managed fleet, not a live public feed. Clients get their own slice in real time, plus the raw alert stream and incident history in a shared channel.
Checkout API — northbank/ledger
Synthetic check on POST /v2/checkout failed from 2 of 4 regions. 5xx rate 6.2%, threshold 1.0%.
Primary on-call acknowledged in 31s. Status thread opened in the shared channel, incident commander named.
Rolled back web@4.18.2, drained us-east-1, raised the pool ceiling 40 → 120. 5xx rate back under 0.05%.
Queries consolidated behind one prepared statement, released as 4.18.3. Postmortem published within 48 hours.
Root cause · Connection pool ceiling of 40 was reached under a promotional traffic spike; release 4.18.2 had added two queries per request path.
A P1 pages a human directly. No acknowledgement in 10 minutes escalates to the secondary on-call, and to the engineering director at 20. The rotation is six SREs on one-week shifts, paid as a line item rather than absorbed into salary.
Contracts get read carefully for a reason. If something here is not the answer you need, ask us directly — we would rather have the awkward conversation before the agreement than after the incident.
Yes — roughly a third of the systems we maintain were written by someone else, and several by someone who is no longer answering email. We start with a two-week takeover: a codebase and infrastructure audit, a dependency and CVE inventory, runbook and access handover, and a written risk register. The SLA starts once that audit is signed off, typically 10 to 15 business days in. If the system cannot be safely supported in its current state, we tell you what has to change first instead of selling you a plan we cannot honour.
Send us the stack, the incident history and the thing you are quietly worried about. You will get a scoped plan and a response-time commitment within two business days — no discovery call required first.