Skip to content
PeakOps
PeakOps HPC ServicesFully remote · annual

We run your on-prem HPC.
You run the science.

A fully remote managed service for the cluster in your own data center, on an annual subscription. Our HPC engineers run and support it remotely — Slurm-native, attached to what you already have, air-gapped capable, on a real SLA. You open tickets; we handle them. A team on tap, not another tool to license.

peakops · cluster health
CPU load72%
Memory64%
Fabric41%
NODESTATELOAD
gpu-a100-01 alloc94%
gpu-a100-02 alloc88%
cpu-hi-07 idle6%
cpu-hi-08 mix51%
cpu-hi-09 drain0%

The visibility and control we give you as part of the service

The problem

On-prem HPC is hard, and specialists are rare.

A cluster is a living system. Keeping it healthy, fair and fully utilized is a full-time discipline — one most teams can’t justify hiring for. So it degrades quietly, and the science slows down. That’s the job we take off your plate.

Nodes drain silently; nobody notices until a run fails.

Slurm config drift — QoS and fair-share nobody fully understands.

The one person who knew the cluster left the group.

GPUs sit idle while jobs pile up in the wrong partition.

Security needs it air-gapped, so cloud tooling is off the table.

What the service covers

Day-2 management, built for how HPC actually works.

The service includes a management console — the same panes our engineers use to run your cluster remotely, shared with your team for full visibility. You’re not buying software; you’re getting the people and the tooling together.

Cluster health

We watch the whole cluster, so you don’t have to.

Live node states, load, memory and interconnect pressure across every partition. As part of the service we surface drains, downs and hot spots — and act on them — before they turn into failed jobs.

  • Per-node state & load
  • Fabric / memory pressure
  • Live across all partitions
peakops · cluster health
CPU load72%
Memory64%
Fabric41%
NODESTATELOAD
gpu-a100-01 alloc94%
gpu-a100-02 alloc88%
cpu-hi-07 idle6%
cpu-hi-08 mix51%
cpu-hi-09 drain0%
Scheduling & QoS

Slurm tuned by people who reason about it daily.

We inspect the queue, tune partitions and QoS, and rebalance fair-share on your behalf — no more hand-editing configs on the head node under pressure. Our console speaks scontrol, squeue and sacctmgr natively.

  • Queue & partition control
  • QoS limits & priority weights
  • Fair-share, made visible
squeue · partition scheduler
JOBIDUSERQOSSTTIME
48213a.yilmazhighR02:14
48219m.demirnormalR00:41
48224lab-cfdnormalPD
48231s.kayalowPD
2 running · 2 pendingfair-share balanced
Alerts & incidents

We know before your researchers do.

We watch node health, partition saturation and accounting lag, and route what matters to your named on-call engineer. Fewer surprises, faster recovery, a clear paper trail — that’s the SLA at work.

  • Node down / drain alerts
  • Saturation & queue backlog
  • Routed to named on-call
peakops · alerts & incidents
node gpu-a100-05 went DOWN
NHC · GPU ECC errors
2m
partition gpu at 96% for 40m
sustained saturation
11m
slurmdbd rollup lag 8m
accounting
26m
node cpu-hi-09 drained
reason: kernel update
1h
1 critical · 2 warningrouted → on-call engineer
Node lifecycle

Nodes drained, resumed and retired cleanly.

Every Slurm node state — idle, alloc, mix, drain, down, resv, power — with the reason string attached, managed from one console. We handle maintenance windows and node rotations so nothing goes dark unnoticed.

  • Full state visibility + reason
  • Drain / resume / power actions
  • Maintenance windows
scontrol · node lifecycle
gpu-a100-05
Reason: GPU ECC errors
draining
idle
alloc
drain
down
resume
drain
resume
power_down
Users & accounting

Accounts, users and usage — reconciled for you.

We manage Slurm accounts and users, map them to your groups, and track service-unit consumption per lab. You get a clear answer to who is using the cluster, and how fairly — ready for reporting.

  • Account & user management
  • Per-group SU accounting
  • Fair-share allocation
sacctmgr · accounts & fair-share
physics-lab12 users4.2M SU34%
genomics8 users3.1M SU26%
cfd-group5 users2.6M SU22%
shared21 users2.0M SU18%
Capacity planning

We tell you when you’ll run out of headroom.

We trend sustained utilization and project when you’ll saturate — so procurement is a planned conversation, not a fire drill. When it’s time for the next rack, you have the evidence to justify it.

  • Utilization trends
  • Saturation forecasting
  • Procurement evidence
peakops · capacity forecast
Sustained utilization
91%
expand in ~6 wk
Q1Q2Q3Q4 (proj.)
And the rest of the engagement

Everything an operator reaches for — handled for you.

Job & queue insight

Per-job history from sacct — runtime, wait, exit codes and TRES — so we can see what’s actually starving the queue, and fix it.

Reservations & maintenance

We schedule maintenance windows and reservations without stranding running jobs or forgetting to lift them afterwards.

Identity integration

We map to your existing LDAP / FreeIPA and Slurm operator accounts. No parallel user directory to keep in sync.

Reporting & exports

Utilization, fair-share and accounting reports exported for grant reporting, chargeback and procurement cases.

Role-based access

Operators, viewers and admins see the right surface. Read-only dashboards for PIs; write actions gated to our team and yours.

Audit trail

Every write action — QoS change, node drain, account edit — is logged with who, what and when. Air-gapped friendly.

How we engage

We work inside your data center.
Fully offline where it counts.

The service runs on your infrastructure and talks only to your cluster’s control plane — slurmctld, slurmdbd, the nodes. No outbound internet required. For secure and classified environments, we operate completely air-gapped.

01
Attach

We stand up a management host on your network and connect it to your existing Slurm through scontrol / sacctmgr and slurmrestd where enabled. No slurmctld replacement, no cluster reinstall — we work with what you run.

02
Air-gap

For secure and classified sites we deliver everything offline — container images and dependencies vendored, zero phone-home. The service runs entirely on your control plane, and so do we.

03
Run

From there it’s our engineers on an SLA: proactive tuning, incident response and upgrades staged on your schedule with rollback if a check fails. Nothing changes behind your back.

Air-gapped capableOn your hardwareNo data leaves siteNamed engineers
engagement · on-prem
peakops ── slurmctld   (control)
        ├─ slurmdbd     (accounting)
        └─ compute[01-32]

host:     management VM on your LAN
talks:    scontrol · sacctmgr · slurmrestd
network:  site-internal only
egress:   none  (air-gapped)
model:    managed service, annual
delivery: fully remote / online
auth:     your LDAP / FreeIPA
support:  named engineers, SLA

Compatible with Slurm 21.08 and newer. slurmrestd used where present; CLI-over-SSH is the reliable baseline everywhere else.

Service packages

An annual subscription, sized by ticket allowance.

One clean annual price per tier — all delivered fully remote / online. Tiers are differentiated by how many support tickets you can open each year and how fast we respond. You open tickets; our engineers handle them. On-site engagements are available as a custom add-on.

StandardFoundation
€1,900/ year
20
support tickets / yearFully remote

Next-business-day response

Fully remote day-2 support for your cluster. You open tickets; our engineers handle them.

  • Slurm scheduling & QoS / fair-share
  • Cluster health monitoring
  • User & job support
  • Next-business-day response
ProMost teams
€4,500/ year
60
support tickets / yearFully remote

Same-business-day response

More allowance, faster SLA and proactive care from engineers who know your cluster.

  • Everything in Standard
  • Same-business-day response
  • Proactive health & capacity reviews
  • Upgrade guidance + quarterly review
EnterpriseMission-critical
Customscoped to site
priority / unlimited ticketsFully remote

Fastest SLA + incident hotline

Deep engagement for large or classified sites — dedicated engineer and the fastest SLA.

  • Everything in Pro
  • Dedicated engineer & fastest SLA
  • Incident hotline
  • Air-gapped support

Research & academic pricing: 50% off any tier. All tiers delivered fully remote / online; on-site engagements available as a custom add-on. Multi-year and multi-cluster engagements scoped on request.

Technical FAQ

The questions your HPC lead will ask.

Straight answers on Slurm, air-gapped delivery and how we run alongside your team — no hand-waving.

Through the interfaces Slurm already exposes: scontrol, squeue, sinfo, sacct and sacctmgr over SSH to the head node, plus slurmrestd where it’s enabled. Historical usage comes from sacct / sshare or the slurmdbd accounting database. We operate as a Slurm operator/admin account for write actions — no agent on your compute nodes.

Contact

Let’s talk about
your cluster.

Standing up on-prem HPC, wrangling Slurm, or bursting to AWS? Tell us what you’re running. A PeakOps engineer will get back to you — not a bot.

hello@peakops.co

No newsletters. We only use this to reply to you.