We run your on-prem HPC.
You run the science.
A fully remote managed service for the cluster in your own data center, on an annual subscription. Our HPC engineers run and support it remotely — Slurm-native, attached to what you already have, air-gapped capable, on a real SLA. You open tickets; we handle them. A team on tap, not another tool to license.
| NODE | STATE | LOAD |
|---|---|---|
| gpu-a100-01 | alloc | 94% |
| gpu-a100-02 | alloc | 88% |
| cpu-hi-07 | idle | 6% |
| cpu-hi-08 | mix | 51% |
| cpu-hi-09 | drain | 0% |
The visibility and control we give you as part of the service
On-prem HPC is hard, and specialists are rare.
A cluster is a living system. Keeping it healthy, fair and fully utilized is a full-time discipline — one most teams can’t justify hiring for. So it degrades quietly, and the science slows down. That’s the job we take off your plate.
Nodes drain silently; nobody notices until a run fails.
Slurm config drift — QoS and fair-share nobody fully understands.
The one person who knew the cluster left the group.
GPUs sit idle while jobs pile up in the wrong partition.
Security needs it air-gapped, so cloud tooling is off the table.
Day-2 management, built for how HPC actually works.
The service includes a management console — the same panes our engineers use to run your cluster remotely, shared with your team for full visibility. You’re not buying software; you’re getting the people and the tooling together.
We watch the whole cluster, so you don’t have to.
Live node states, load, memory and interconnect pressure across every partition. As part of the service we surface drains, downs and hot spots — and act on them — before they turn into failed jobs.
- Per-node state & load
- Fabric / memory pressure
- Live across all partitions
| NODE | STATE | LOAD |
|---|---|---|
| gpu-a100-01 | alloc | 94% |
| gpu-a100-02 | alloc | 88% |
| cpu-hi-07 | idle | 6% |
| cpu-hi-08 | mix | 51% |
| cpu-hi-09 | drain | 0% |
Slurm tuned by people who reason about it daily.
We inspect the queue, tune partitions and QoS, and rebalance fair-share on your behalf — no more hand-editing configs on the head node under pressure. Our console speaks scontrol, squeue and sacctmgr natively.
- Queue & partition control
- QoS limits & priority weights
- Fair-share, made visible
| JOBID | USER | QOS | ST | TIME |
|---|---|---|---|---|
| 48213 | a.yilmaz | high | R | 02:14 |
| 48219 | m.demir | normal | R | 00:41 |
| 48224 | lab-cfd | normal | PD | — |
| 48231 | s.kaya | low | PD | — |
We know before your researchers do.
We watch node health, partition saturation and accounting lag, and route what matters to your named on-call engineer. Fewer surprises, faster recovery, a clear paper trail — that’s the SLA at work.
- Node down / drain alerts
- Saturation & queue backlog
- Routed to named on-call
Nodes drained, resumed and retired cleanly.
Every Slurm node state — idle, alloc, mix, drain, down, resv, power — with the reason string attached, managed from one console. We handle maintenance windows and node rotations so nothing goes dark unnoticed.
- Full state visibility + reason
- Drain / resume / power actions
- Maintenance windows
Accounts, users and usage — reconciled for you.
We manage Slurm accounts and users, map them to your groups, and track service-unit consumption per lab. You get a clear answer to who is using the cluster, and how fairly — ready for reporting.
- Account & user management
- Per-group SU accounting
- Fair-share allocation
| physics-lab | 12 users | 4.2M SU | 34% |
| genomics | 8 users | 3.1M SU | 26% |
| cfd-group | 5 users | 2.6M SU | 22% |
| shared | 21 users | 2.0M SU | 18% |
We tell you when you’ll run out of headroom.
We trend sustained utilization and project when you’ll saturate — so procurement is a planned conversation, not a fire drill. When it’s time for the next rack, you have the evidence to justify it.
- Utilization trends
- Saturation forecasting
- Procurement evidence
Everything an operator reaches for — handled for you.
Job & queue insight
Per-job history from sacct — runtime, wait, exit codes and TRES — so we can see what’s actually starving the queue, and fix it.
Reservations & maintenance
We schedule maintenance windows and reservations without stranding running jobs or forgetting to lift them afterwards.
Identity integration
We map to your existing LDAP / FreeIPA and Slurm operator accounts. No parallel user directory to keep in sync.
Reporting & exports
Utilization, fair-share and accounting reports exported for grant reporting, chargeback and procurement cases.
Role-based access
Operators, viewers and admins see the right surface. Read-only dashboards for PIs; write actions gated to our team and yours.
Audit trail
Every write action — QoS change, node drain, account edit — is logged with who, what and when. Air-gapped friendly.
We work inside your data center.
Fully offline where it counts.
The service runs on your infrastructure and talks only to your cluster’s control plane — slurmctld, slurmdbd, the nodes. No outbound internet required. For secure and classified environments, we operate completely air-gapped.
We stand up a management host on your network and connect it to your existing Slurm through scontrol / sacctmgr and slurmrestd where enabled. No slurmctld replacement, no cluster reinstall — we work with what you run.
For secure and classified sites we deliver everything offline — container images and dependencies vendored, zero phone-home. The service runs entirely on your control plane, and so do we.
From there it’s our engineers on an SLA: proactive tuning, incident response and upgrades staged on your schedule with rollback if a check fails. Nothing changes behind your back.
peakops ── slurmctld (control)
├─ slurmdbd (accounting)
└─ compute[01-32]
host: management VM on your LAN
talks: scontrol · sacctmgr · slurmrestd
network: site-internal only
egress: none (air-gapped)
model: managed service, annual
delivery: fully remote / online
auth: your LDAP / FreeIPA
support: named engineers, SLACompatible with Slurm 21.08 and newer. slurmrestd used where present; CLI-over-SSH is the reliable baseline everywhere else.
An annual subscription, sized by ticket allowance.
One clean annual price per tier — all delivered fully remote / online. Tiers are differentiated by how many support tickets you can open each year and how fast we respond. You open tickets; our engineers handle them. On-site engagements are available as a custom add-on.
Next-business-day response
Fully remote day-2 support for your cluster. You open tickets; our engineers handle them.
- Slurm scheduling & QoS / fair-share
- Cluster health monitoring
- User & job support
- Next-business-day response
Same-business-day response
More allowance, faster SLA and proactive care from engineers who know your cluster.
- Everything in Standard
- Same-business-day response
- Proactive health & capacity reviews
- Upgrade guidance + quarterly review
Fastest SLA + incident hotline
Deep engagement for large or classified sites — dedicated engineer and the fastest SLA.
- Everything in Pro
- Dedicated engineer & fastest SLA
- Incident hotline
- Air-gapped support
Research & academic pricing: 50% off any tier. All tiers delivered fully remote / online; on-site engagements available as a custom add-on. Multi-year and multi-cluster engagements scoped on request.
The questions your HPC lead will ask.
Straight answers on Slurm, air-gapped delivery and how we run alongside your team — no hand-waving.
Through the interfaces Slurm already exposes: scontrol, squeue, sinfo, sacct and sacctmgr over SSH to the head node, plus slurmrestd where it’s enabled. Historical usage comes from sacct / sshare or the slurmdbd accounting database. We operate as a Slurm operator/admin account for write actions — no agent on your compute nodes.
No. We run from a separate management host and attach to your existing cluster. slurmctld and slurmdbd stay exactly as they are. If you ever end the engagement, your cluster keeps running unchanged — there is no lock-in, because we manage standard Slurm on your standard hardware.
Genuinely air-gapped. Everything is delivered offline with all container images and dependencies vendored — no phone-home, no telemetry, no runtime downloads. The only network traffic is between the management host and your control plane on your internal network. The honest caveat: first install and every update arrive as a physically transferred bundle, and our team works within your site’s access controls.
Slurm 21.08 and newer. The CLI tools (scontrol, sacct, sacctmgr) are stable across that range, so they’re the reliable baseline. slurmrestd is used where present, but since it must be explicitly built and its API version varies between releases, we never depend on it being available.
Yes — we read and edit the real Slurm knobs. QoS via sacctmgr (MaxJobs, MaxTRESPerUser, GrpTRES, priority weights), the multifactor priority weights in slurm.conf, and the fair-share association tree with its decay half-life. We make what’s already configured visible and tunable, rather than imposing a parallel model on top of your cluster.
The full set Slurm reports — idle, allocated, mixed, completing, drain/draining/drained, down, reserved, maintenance and power states — each with its reason string. Operator actions map straight to scontrol update: drain, resume, down and power_down/up, with the reason recorded and visible in the audit trail we keep for you.
Both, sold as one service. You get named HPC engineers on an SLA who run day-2 operations remotely, plus the management console they use — shared with your team for visibility. It’s an annual subscription priced by support-ticket allowance and response SLA, delivered fully remote / online — not a software license. Upgrades and the console come with the service; there’s no separate seat or license to buy. On-site engagements are available as a custom add-on.
Let’s talk about
your cluster.
Standing up on-prem HPC, wrangling Slurm, or bursting to AWS? Tell us what you’re running. A PeakOps engineer will get back to you — not a bot.
hello@peakops.co