Skip to main content
You are here: Releases Notes

2.5.1 Release Notes

Summary -

Overview - Release V2.5.1 is a targeted patch release focused on GPU-sharing (HAIShare) reliability and Admin Panel visibility across the hosted·ai | GPUaaS platform for service providers Platform. Key fixes address orphaned GPU pool and worker records that blocked capacity reclamation, tenant fairness and startup reliability on shared GPU cards, mixture-of-experts model support under GPU sharing and load-balancing defaults for newly created GPU-sharing pools.

Bug fixes

1. GPU Sharing (HAIShare)

Ticket

Description

Priority

Business impact / value to end users

HAI-13729

Deleting a team or its GPU workloads could leave phantom worker entries behind in the platform's resource manager, permanently blocking the GPU pool — and its GPUs — from ever being removed, even though the workload itself was long gone.

Critical

Administrators can now reliably delete idle GPU pools and reclaim their GPUs after workloads and teams are removed, without needing manual backend intervention to clear stuck records.

HAI-13753

GPU-sharing teardown command could report a failure when removing a tenant that had already been released as part of the same cleanup, and the scheduler's own confirmation message was not visible in standard logs.

High

Makes GPU pool and instance teardown more resilient by ensuring normal cleanup steps are no longer misreported as failures, reducing false alerts and stale pool state for administrators.

HAI-13743

On a shared GPU pool under load, a small number of tenants could receive almost no GPU service — a fraction of their expected throughput — while the rest of the pool had spare, idle capacity available.

High

Tenants sharing a GPU pool now receive a fair, comparable share of GPU throughput under load, preventing a workload from being effectively starved while capacity sits unused elsewhere in the pool.

HAI-13530

Running mixture-of-experts models (such as Qwen3 MoE, DeepSeek, or Mixtral) on GPU-sharing instances could fail outright with a CUDA error, while the same underlying issue also quietly slowed down every model sharing that code path.

High

Mixture-of-experts models — an increasingly common architecture for modern LLMs — now run reliably on shared GPU instances, and all models benefit from the removal of an unnecessary performance-limiting step.

HAI-13529

On a fully-utilized shared GPU card, a new workload attempting to start could fail immediately with an out-of-memory error, even though the workload's own view of the card showed no memory in use, leaving no clear indication of the real cause.

Blocker

New workloads placed on busy shared GPU cards now start successfully instead of failing at launch, removing a hard blocker for customers deploying on GPU-sharing instances under load.

HAI-13069

When an administrator created a new GPU-sharing pool without explicitly choosing a load-balancing policy, the platform always forced it onto round-robin scheduling, silently overriding the GPU-sharing scheduler's own configured default.

Critical

New GPU-sharing pools created without an explicit policy now correctly follow the scheduler's intended default behaviour, avoiding unintentionally under-optimized workload placement.

HAI-13954

If the admin panel restarted while a pod provision or deprovision recipe was mid-flight, the recipe replayed and completed successfully on restart, but its completion handler was lost — leaving the pod reporting "Provisioning" or "Deprovisioning" indefinitely with nothing to clear it.

High

Pod provisioning and deprovisioning now complete correctly even if the panel restarts mid-operation, preventing pods from getting stuck "in progress" indefinitely and previously requiring manual database correction to clear.

HAI-13979

The automatic retry job for failed pod teardowns filtered on a status value that nothing in the system ever wrote, so it matched zero rows on every run — failed deprovisions were silently never retried and never escalated, even after months of accumulation.

High

Failed pod teardowns are now automatically retried and escalated for manual cleanup after repeated failures, reclaiming GPU capacity that was previously left stuck and held away from every buyer indefinitely.

HAI-13981

The pod deprovision recipe treated a "GPU is not assigned to team" response — the expected outcome when a GPU hold was already released — as a fatal error, so a teardown could never finish for a pod whose haiDra hold was already gone, even after repeated retries.

High

Pod teardown now completes successfully in this already-released scenario instead of failing repeatedly and requiring manual cleanup, freeing up GPU capacity that was being incorrectly withheld from buyers.

HAI-13996

The deprovision recipe always issued a second, redundant per-GPU unsubscribe call after the primary release had already succeeded; on older haiDra versions (below 8.6.32) that fallback call was not correctly scoped to the specific GPU and could release a different, still-running pod's hold in the same pool.

High

Eliminates a rare case on older haiDra versions where tearing down one pod could inadvertently release the GPU hold of another active pod, preventing a running workload from being unexpectedly reassigned or interrupted.

2. Admin Panel

Ticket

Description

Priority

Business impact / value to end users

HAI-13077

"Teams using this pool" tab on the GPUaaS pool details page in the Admin Panel showed no teams for any pool, even when teams had active GPU workloads running on it.

High

Administrators now see an accurate list of which teams are actively using a given GPU pool, supporting capacity planning and customer support without needing to query the database directly.

3. MESH

Ticket

Description

Priority

Business impact / value to end users

MHA-1009

My Subscriptions only loaded the 50 most recent rows and filtered client-side, so a buyer with churned subscriptions could have their one live subscription hidden from view.

Normal

Buyers with a long subscription history now reliably see their active subscription instead of it being hidden behind older, churned entries.

MHA-1001

The haiDra share-to-vGPU conversion double-scaled above sharing ratio 2, so a fully-held ratio-4 pool reported 4x its actual held capacity.

Normal

GPU pool capacity is now reported accurately at higher sharing ratios, preventing pools from appearing far more utilized than they really are.

MHA-1004

A stale haiDra timestamp kept consumed_vgpus artificially high for roughly 14 minutes after a vGPU was actually freed.

Normal

Freed GPU capacity becomes visible to buyers almost immediately instead of appearing unavailable for up to 14 minutes after release.

MHA-971

The buyer panel's use of omitempty dropped a genuine available_vgpus=0 response, causing the UI to fabricate a value of 1 or 4 instead of showing zero availability.

High

Buyers now see accurate GPU availability, including when a pool is genuinely full, instead of being shown fabricated open capacity.

MHA-1000

Mesh pool capacity is now resolved live when the pool list is built, rather than from stale cached values.

Normal

Buyers are never offered a mesh pool they can't actually deploy into, avoiding failed deployments caused by outdated capacity data.

MHA-1002

The user-panel's copy of a clone pool never received total_vgpus/consumed_vgpus, so physicalKnown was always false.

Normal

Clone pools now correctly report physical GPU capacity in the user panel instead of always appearing as unknown.

MHA-1006

A delete request accepted during a panel restart could leave a Running pod row with no haiDra hold, and nothing retried it — withholding the pool from every buyer.

Normal

Pools are no longer incorrectly withheld from all buyers due to an orphaned pod record left behind by a delete that raced a panel restart.

MHA-1007

pod_recon reconciled active_deployments against a raw POD count, incorrectly healing a multi-vGPU pod's counter downward and oversubscribing the pool.

Normal

Multi-vGPU pods are now counted correctly during reconciliation, preventing pools from being oversold beyond their real capacity.

MHA-1008

Pod provision and deprovision recipes were not ordered per worker, so a deprovision that finished first could leave an untracked pod still running on the GPU.

High

Provision and deprovision operations on the same worker are now properly sequenced, eliminating orphaned pods left running on a GPU after teardown.

Upgrade instructions

This is only an incremental upgrade and it is to be applied on 2.5.0

Panel upgrade

apt update apt install --only-upgrade -y hostedai-adminpanel-api apt install --only-upgrade -y hostedai-adminpanel-ui apt install --only-upgrade -y hostedai-userpanel-api apt install --only-upgrade -y hostedai-userpanel-ui

Recipe upgrade

Update recipes using deploy utility

./deploy --recipes

HaiDra upgrade

apt update apt install --only-upgrade -y haidra