Threadripper Halo Station: 96 Cores, 576GB [2026]
Tuesday, September 08, 2026AMD Threadripper Halo Station: 96 Core plus MI350P 576GB GPU and 2TB System Per VideoCardz Sept 4 [2026 Full Guide]
By Indie Kings | September 8, 2026 | Updated September 8, 2026 | 19 min read

Photo: Pexels free use.
On September 4, 2026, VideoCardz reported AMD Threadripper Halo Station packs a 96 core CPU and Instinct MI350P GPUs with up to 576GB GPU memory and 2TB system memory. AMD showed the liquid cooled deskside system at IFA 2026 with a 96 core Threadripper PRO CPU, two MI350P cards for 288GB in the demo unit, and a path to four cards for 576GB HBM3E. The stated goal is local trillion parameter class AI without cloud. Price is unknown and no ship date is confirmed. This full guide covers confirmed specs, workstation versus desktop differences, memory math for trillion parameter models, power and cooling reality, who should buy versus who should skip, ROCm software caveats, price modeling from component street, setup planning, and what remains unconfirmed. For live workstation plus GPU bundle prices see our Bundle Tracker.
Table of Contents
- 1. What VideoCardz Reported Sept 4 at IFA 2026
- 2. Full Spec Sheet: CPU, GPU, Memory, IO
- 3. Workstation vs Desktop: Why This Is Not a Gaming PC
- 4. Memory Math: 576GB HBM3E plus 2TB DDR5 for Trillion Params
- 5. Power Cooling Noise and Room Planning
- 6. Price Unknown: Component Model and 100K Plus Logic
- 7. Who It Fits: Labs, Studios, Local AI Teams
- 8. Software Stack: ROCm, Frameworks, Validation Gaps
- 9. Alternatives: DGX Station, Cloud Rent, Smaller Workstations
- 10. Verdict and What To Watch Next
1. What VideoCardz Reported Sept 4 at IFA 2026
VideoCardz published the Halo Station report on September 4, 2026, summarizing AMD IFA 2026 keynote details. Core facts: Threadripper Halo Station with a 96 core CPU, Instinct MI350P GPUs, up to 576GB GPU memory and 2TB system memory. The demo configuration used two MI350P PCIe cards for 288GB HBM3E. The planned full configuration supports four GPUs for 576GB HBM3E. Each MI350P carries 144GB HBM3E at up to 4TB per s bandwidth. The CPU maps to Ryzen Threadripper PRO 9995WX, the only current PRO model with 96 Zen 5 cores and 192 threads, 8 channel DDR5, 128 PCIe 5.0 lanes, and up to 5.4 GHz boost. Cooling is full liquid for CPU and accelerators in a deskside tower. AMD frames it as the most powerful workstation in the world for local models above one trillion parameters. No price, no firm ship date, and no confirmed OEM partners were named on stage per companion coverage. Jack Huynh, AMD SVP and GM of Computing and Graphics, presented it as a new class that brings supercomputer class compute to individual users and developers for local train, fine tune, and agentic runs. That is the full reported base. Everything else in this guide is planning context around those points. See workstation bundle value in our Bundle Tracker.
Why IFA timing matters is straightforward. IFA 2026 opened in Berlin in early September, with AMD holding the opening keynote. TechPowerUp staff on the ground, plus Tomshardware, TechSpot, ServeTheHome, and Igorslab, corroborated the same configuration within 48 hours. That fast corroboration raises confidence that the demo unit specs are real, while price and date remain open because AMD did not post them. Prototype wording also matters. AMD calls Halo Station a prototype system shown for the first time at IFA 2026. Prototype means chassis, loop, firmware, and board layout can still change before retail. It also means benchmarks cited by AMD are vendor directed until third parties test a shipping unit. Treat the trillion parameter claim as capacity math plus vendor workload tuning, not as plug and play proof for every framework. Our job below is to separate confirmed hardware from unconfirmed cost, date, performance, and software readiness.
2. Full Spec Sheet: CPU, GPU, Memory, IO
CPU is Ryzen Threadripper PRO 9995WX, Zen 5 Shimada Peak flagship for workstations. It packs 96 cores and 192 threads, boost up to 5.4 GHz, 350W TDP class, 8 channel DDR5 memory controller, and 128 usable PCIe 5.0 lanes. Eight channels are what enable up to 2TB DDR5 system memory with high capacity RDIMMs. That system RAM holds datasets, CPU preprocess, caches, and overflow for models that do not fully fit in HBM3E. One hundred twenty eight PCIe lanes are what make two to four full x16 accelerators plus fast NVMe plus 10GbE plus USB4 possible without lane sharing pain. This is not a repurposed gaming chip. It is a workstation CPU with PRO manageability, ECC support paths, and multi day load validation. For mixed CPU plus GPU AI work, the 96 cores feed data loaders, tokenizers, and CPU side graph work while GPUs stay fed. For classic workstation work like compile, simulation preprocess, and video encode, the same cores crush threaded tasks that stall 16 and 24 core desktops. Pair with enterprise NVMe RAID for checkpoint speed. Compare DDR5 RDIMM bundle cost in our Bundle Tracker.
Accelerators are AMD Instinct MI350P PCIe cards based on CDNA 4. Each card has 128 Compute Units, 144GB HBM3E, up to 4TB per s bandwidth, PCIe 5.0 x16 host link, and up to 600W TBP. Two cards in the IFA demo equal 288GB accelerator memory and up to 8TB per s combined. Four cards equal 576GB and up to 16TB per s combined. HBM3E bandwidth near 4TB per s per card is about 14 times the stream bandwidth of LPDDR5X class unified mini PCs, which is why large model token rates and large batch fine tunes prefer HBM even when unified pools look large on paper. The PCIe card form factor matters for serviceability versus SXM style baseboards. PCIe cards can be swapped, upgraded, and cooled in a tower loop, while SXM needs a server chassis. The tradeoff is link speed. PCIe 5.0 x16 between cards is far slower than server scale Infinity Fabric, so multi GPU collectives need careful tuning. Combined memory with system RAM reaches near 2.6TB when 2TB DDR5 plus 576GB HBM3E are counted together, though only HBM is fast enough for hot weights. Use system RAM for data staging, HBM for active model.
3. Workstation vs Desktop: Why This Is Not a Gaming PC
A gaming desktop optimizes for low latency frames, bursty clocks, 16 to 32GB VRAM, 1 to 2TB NVMe, 850 to 1200W PSU, and air or AIO cooling in a mid tower. Halo Station optimizes for sustained throughput, multi kilowatt continuous load, hundreds of gigabytes of HBM residency, terabytes of DDR5, ECC style data integrity, remote manageability, and 24/7 liquid cooling in a deskside chassis with server grade airflow. The parts overlap in name only. Threadripper PRO needs a WRX90 class board with 8 channel trace layout, beefy VRM with active cooling, and BMC style management, not a B650 gaming board. MI350P needs full length triple slot PCIe clearance, auxiliary power at data center levels, and liquid quick disconnects, not a single 12V 2x6 gaming plug. System validation targets weeks of checkpointed training without silent corruption, not a three hour gaming session. Drivers split too. Gaming uses GeForce or Radeon Adrenalin tuned for frame pacing. Halo uses ROCm plus Instinct data center drivers tuned for math correctness and multi GPU collectives. You can technically launch a game on workstation GPUs, but you pay massive cost per frame and lose gaming driver polish. Do not buy Halo for high refresh gaming. Buy a 9800X3D plus RTX 5080 class desktop for one tenth the money and get better fps. Find gaming value in our Bundle Tracker.
Ownership model also differs. Desktops assume a home office with standard 15 amp circuits, moderate noise tolerance, and user service with a screwdriver. Halo assumes a lab or studio with dedicated power, ventilation, and IT handling. Two MI350P cards can draw up to 1200W combined plus 350W CPU before memory, drives, fans, pumps, and network. A four GPU build can exceed 2.4kW for accelerators alone, pushing wall draw past 3kW under load. That needs more than one standard room circuit, careful breaker planning, and cooling that dumps server heat continuously. Noise is pump plus radiator fan noise under constant load, not idle quiet with gaming bursts. Physical size is deskside tower, not mini PC. Weight is server class. Warranty expects next business day onsite or depot with data handling terms, not consumer RMA mail in. Software expects Linux first, Windows second if at all for full ROCm features. If your team has no Linux admin and no dedicated power, you are not the buyer yet. Rent cloud or buy a smaller 128GB unified mini workstation first.
| Trait | Gaming desktop | Halo Station workstation |
|---|---|---|
| Goal | Low latency frames | Sustained AI throughput for days |
| CPU | 8 to 24 cores, bursty boost | 96 cores, 192 threads, ECC paths |
| GPU memory | 16 to 32GB GDDR7 | 288GB demo, 576GB max HBM3E |
| System RAM | 32 to 96GB DDR5 | Up to 2TB DDR5 RDIMM |
| Power | 850 to 1200W PSU | Multi kW continuous, dedicated circuits |
| Cooling | Air or 360mm AIO | Full liquid CPU plus GPUs, 24/7 loop |
| Software | Game drivers, Windows first | ROCm, Linux first, multi GPU tuned |
| Price band | 1500 to 4000 dollars | 100K plus speculative, unconfirmed |
4. Memory Math: 576GB HBM3E plus 2TB DDR5 for Trillion Params
Model memory math decides whether trillion parameter claims make sense. At 4 bit quantization, each parameter needs about 0.5 bytes plus overhead for activations, KV cache, and runtime buffers. One trillion parameters at 4 bit therefore needs roughly 500GB for weights alone before overhead. A 576GB HBM3E pool across four MI350P cards can hold those weights on device with limited headroom left for context, which matches AMD framing that a 4 bit quantized trillion parameter class model can fit within GPU memory. At 8 bit, the same model needs near 1TB for weights, which exceeds 576GB and forces offload to system DDR5 or multi node sharding. At 16 bit, needs double again. Practical local work will therefore use 4 bit quants, mixture of experts offload tricks, and short to medium context to stay resident. The two card demo with 288GB HBM3E fits roughly 500B plus parameter class at 4 bit with room for context, or 200B to 400B dense models comfortably with long context and large batches. That already covers most open models labs actually run today. Four cards extend the ceiling to trillion class demos, large mixture of experts with many active experts, and multi model agentic pipelines where two models stay resident at once. System DDR5 at up to 2TB does not replace HBM speed, but it stages datasets, holds CPU offloaded experts, caches checkpoints, and absorbs burst context without crashing. Think of HBM as hot stage, DDR5 as warm backstage. Track high capacity DDR5 pricing in our Bundle Tracker.
Bandwidth math matters as much as capacity. Each MI350P peaks near 4TB per s HBM bandwidth. Four cards peak near 16TB per s aggregate if perfectly parallel, though real collectives and PCIe hops lower sustained rates. Token generation speed scales with how fast weights stream to compute each step. HBM at terabytes per second sustains far higher tokens per second than DDR5 at hundreds of GB per s or LPDDR5X unified pools at similar hundreds of GB per s. That is why Halo can claim local trillion parameter runs while 128GB unified mini PCs target 120B class models. Both have large pools on paper, but HBM feeds active weights an order of magnitude faster. Context length also eats memory fast. A 1M token context with large hidden size can add tens to hundreds of GB for KV cache depending on model and batch. Long context plus large batch is what pushes even 576GB builds to offload. Labs should size by worst case context times batch, not by weight size alone. Log resident set plus KV cache plus activation peaks during trial runs before promising a workload fits.
5. Power Cooling Noise and Room Planning
Power is the hidden spec that decides room readiness. Threadripper PRO 9995WX sits at 350W TDP with higher peaks under all core AVX style loads. Each MI350P lists up to 600W TBP. Two cards plus CPU therefore approach 1550W for compute alone before DDR5 RDIMMs, NVMe arrays, fans, pumps, network, and PSU losses. Real wall draw for the demo class likely lands near 1800 to 2200W under full AI load. Four cards plus CPU approach 2750W for compute alone, with wall draw likely past 3000W once everything is counted. Standard North American 120V 15 amp circuits deliver 1800W max with 1440W continuous safe guidance. One Halo Station can exceed one room circuit by itself. Plan for multiple dedicated 20 amp circuits, or 208 to 240V lab feeds where available, with a licensed electrician confirming breaker, wire gauge, and outlet rating. Do not daisy chain power strips. Use metered PDUs to log draw per phase. Budget UPS only for graceful checkpoint save and shutdown, not for continued training through outages, because multi kW UPS capacity is costly and heavy. Verify PSU redundancy and rail layout with the final OEM, since prototype power plumbing can change. See PDU plus rack bundle options in our Bundle Tracker.
Cooling is full liquid for a reason. Multi kilowatt heat cannot be handled by workstation fans without extreme noise and throttling. Halo uses a liquid loop covering CPU and accelerators in a deskside tower, with external radiator area and controlled fan curves for sustained operation. Expect continuous pump plus fan noise under load, warm exhaust that heats small rooms fast, and maintenance needs like coolant checks, dust filter cleaning, and quick disconnect inspection. Room HVAC must dump 2 to 3kW continuously during long runs. A 15 square meter closed office will overheat in under an hour without ventilation. Plan for exhaust to hallway or dedicated mini split, intake with filtered cool air, and ambient monitoring that pauses jobs above safe inlet temps. Floor loading and footprint also matter. A fully loaded liquid tower with four GPUs, large radiators, and 2TB RAM can exceed 50kg. Verify floor, desk, door width, and elevator access before delivery. Noise sensitive open offices should place Halo in a server closet with 10GbE to desks, not beside creators.
| Config | Compute power rough | Wall draw est with losses | Room need |
|---|---|---|---|
| 2x MI350P demo plus 96 core | 1200W GPUs plus 350W CPU | 1800 to 2200W | Two dedicated circuits, vented room |
| 4x MI350P max plus 96 core | 2400W GPUs plus 350W CPU | 3000W plus | Lab feed, HVAC, server closet |
| Idle plus light prep | 200 to 400W | 250 to 500W | Standard office ok |
| Gaming desktop reference | 400W GPU plus 150W CPU | 600 to 800W | Single circuit, normal cooling |
6. Price Unknown: Component Model and 100K Plus Logic
AMD has not announced official pricing as of September 8, 2026. That single fact should anchor every budget discussion. Outlets have speculated 100,000 to 150,000 dollars for finished configurations, with Tomshardware style component math noting CPU plus 2TB DDR5 plus two accelerators alone exceeds 100,000 dollars at street prices. Treat those figures as unconfirmed estimates, not as MSRP. Why the math still helps is for approval planning. A 96 core Threadripper PRO flagship lists in the five figure range at retail. High capacity DDR5 RDIMMs for 2TB cost tens of thousands at 2026 pricing. Instinct MI350P cards are data center parts with data center pricing, not gaming MSRP. Two cards dominate system cost. Four cards double that dominant line. Add WRX90 class board, multi kW PSU setup, full liquid chassis, enterprise NVMe, 10GbE, assembly, validation, onsite warranty, and margin, and six figures follow quickly. Fully loaded four GPU plus 2TB RAM builds could exceed 150,000 dollars if HBM supply is tight. None of that is official until AMD or an OEM posts a SKU. Use it to reserve budget, not to sign purchase orders. Track component street movement in our Bundle Tracker.
Buying logic should compare own versus rent. Cloud H100 or MI300 class rentals run dollars per GPU hour, with monthly totals in the thousands for always on single GPU and tens of thousands for multi GPU fleets. A lab that needs dedicated always on capacity for 12 to 24 months, values data locality, and runs steady fine tunes can justify six figure capex because payback arrives within a year versus renting equivalent HBM capacity. A team with bursty needs, short pilots, or uncertain model choice should rent first and buy later. Depreciation, power bills, cooling, admin time, and ROCm porting cost all count against own. Resale for Instinct cards is thinner than for gaming GPUs, so assume long hold. Procurement tip: request fixed BOM with RAM, SSD, network, warranty years, and power cords itemized. Ask whether two GPU units can upgrade to four GPUs in field without chassis swap. Confirm whether price includes liquid maintenance kit and spare fans. Get power and noise specs in writing for facilities approval before finance approval, because facilities can veto a workstation that finance already approved.
7. Who It Fits: Labs, Studios, Local AI Teams
Best fit is a research lab or enterprise AI team that needs local trillion parameter class experiments, private data that cannot leave premises, always on agentic pipelines, and staff who already run ROCm or are funded to port. Examples include genomics groups with private cohorts, legal discovery teams with confidential corpora, robotics labs doing large vision language fine tunes, and studios rendering plus generating with large local models where cloud egress fees and queue waits hurt iteration. These buyers value HBM residency, 96 CPU cores for data prep, 2TB DDR5 for staging, and deskside control over job scheduling. They have Linux admins, backup for multi terabyte checkpoints, and power plus cooling already approved for lab gear. They measure value in weeks saved waiting on shared cluster queues, not in fps per dollar. If cloud queues delay every experiment by days, a local Halo Station pays back in schedule, not just in dollars. It also simplifies compliance reviews when data never leaves the building. For these teams, even the two GPU demo class with 288GB HBM3E is a major step beyond 128GB unified mini PCs that top out near 120B class models. See storage plus backup bundles in our Bundle Tracker.
Poor fit is almost everyone else. Solo gamers, streamers, 4K video editors without local LLM needs, small agencies doing cloud SaaS AI, students, and general creators should not buy Halo. A 128GB unified mini PC or a 24GB gaming GPU with cloud burst covers their AI needs for under 5,000 dollars. IT generalists without ROCm experience will struggle with driver plus framework gaps. Offices without dedicated power and ventilation cannot host it safely. Finance teams expecting three year laptop style depreciation will balk at accelerator residual risk. Even well funded startups should pilot on rented HBM first to prove model choice and utilization above 60 percent before committing six figures. The middle ground is real though. A studio that outgrows a mini PC but does not need four GPUs can target the two GPU Halo SKU, keep 1TB DDR5 instead of 2TB, and add GPUs later if the chassis allows field upgrade. That staged path controls risk while keeping the HBM ceiling open.
8. Software Stack: ROCm, Frameworks, Validation Gaps
Hardware without software is a heater. Halo runs on AMD ROCm plus Instinct drivers, with PyTorch, vLLM, Triton style kernels, and container tooling as the practical stack for local LLM work. As of early September 2026, independent validation of that full stack on shipping Halo hardware does not exist because no retail unit is shipping. Precedent from adjacent RTX Spark CUDA gaps shows why buyers must demand proof. New accelerators often launch with great silicon and lagging software, where flagship features work in vendor demos but fail on popular community containers or specific attention kernels. For Halo, ask for validated versions in writing. Which ROCm release, which PyTorch build, which vLLM commit, which quantization kernels for 4 bit trillion class, which multi GPU collective backend across PCIe cards, and which checkpoint formats are tested. Ask for throughput logs, not just capacity claims. Tokens per second at defined context plus batch, time to fine tune a reference LoRA, and multi GPU scaling from two to four cards matter more than fit alone. A box that fits a model but generates at unusable speed is not a solution. Require a live run on your model family before acceptance. Keep cloud fallback during porting. Find Linux plus container bundle resources in our Bundle Tracker.
Operational details decide daily pain. Check Linux distro support first, Ubuntu LTS versus RHEL style, kernel version pins, and secure boot behavior with signed drivers. Check container runtime, registry mirrors, and air gap install paths for private labs. Check monitoring for HBM use, ECC events, thermals, pump status, and per GPU power, with alerts that pause jobs before hardware trips. Check checkpoint strategy, NVMe endurance for constant writes, and backup windows for multi terabyte snapshots. Check user quotas and job scheduling, because a single user can monopolize 576GB HBM and block a team. Check security, user separation, encrypted datasets at rest, and audit logs for compliance. None of this is unique to AMD, but new platforms expose gaps faster. Early buyers should budget two to four weeks for bring up, kernel tuning, and framework pins before promising project timelines. Late buyers will inherit community playbooks and suffer less. That delay tradeoff is the classic early adopter tax.
9. Alternatives: DGX Station, Cloud Rent, Smaller Workstations
Nvidia DGX Station class is the direct alternative for teams that already run CUDA. DGX heritage offers mature software, broad container support, and familiar NCCL collectives, but at data center pricing and often lower HBM ceilings per dollar than Halo four GPU math on paper. Choice hinges on software, not just memory. If your pipelines are CUDA locked with custom kernels and no porting budget, DGX or rented H100 plus B200 capacity wins even if Halo lists more gigabytes. If your stack is PyTorch standard with ROCm compatible kernels and you value local control, Halo becomes credible once validated. Do not decide on gigabytes alone. Decide on validated tokens per second per dollar for your model plus context plus batch, including power and admin cost. Request head to head trials with the same prompt sets and the same quantization before signing. For many labs the rational answer through 2026 is hybrid. Rent Nvidia HBM by the hour for deadline work while piloting Halo for steady local base load. That hedges software risk while capturing locality wins. Compare accelerator rental versus own math in our Bundle Tracker.
Smaller local options cover most creators without Halo money. RTX Spark mini PCs with 128GB unified memory target 120B class models at mini PC power and price. Ryzen AI Max plus workstations with 128 to 192GB unified target similar territory with x86 software familiarity. Single RTX 5090 32GB desktops handle 30B to 70B quants locally with fast iteration for individuals. Cloud APIs handle trillion class reasoning without any capex for teams that can send data off site. The Halo buyer is the narrow slice that needs more than 128GB fast residency, cannot use cloud for data or queue reasons, and can staff and power a multi kW box. If you do not check all three boxes, buy smaller local plus cloud burst and revisit Halo after retail validation. The worst outcome is a six figure deskside that idles because the model choice changed or the team never ports kernels. Utilization below 40 percent destroys ROI versus rent. Measure queue waits and data constraints honestly before committing.
10. Verdict and What To Watch Next
Verdict as of September 8, 2026 is prototype promise, not purchase order. Confirmed: 96 core Threadripper PRO class, two MI350P cards in demo for 288GB HBM3E, path to four for 576GB, up to 2TB DDR5, full liquid deskside, IFA 2026 unveil, trillion parameter local goal. Unconfirmed: price, ship date, OEMs, real third party throughput, four GPU thermals in deskside volume, and full ROCm validation for popular frameworks. Workstation versus desktop verdict is clear. Halo is not a faster gaming PC. It is a deskside data center node for AI teams with power, cooling, Linux talent, and private data needs. Price unknown verdict is also clear. Plan six figures, do not budget to the dollar, and require itemized BOM plus field upgrade path in writing. Most readers should admire, not order. Labs with queue pain and locality mandates should engage AMD for trial terms while keeping cloud capacity. Everyone else should buy 128GB unified mini or 24GB gaming plus cloud and wait for retail proof. Revisit after OEM SKUs post with power, noise, warranty, and validated ROCm versions. Track OEM listings in our Bundle Tracker.
Watch next in order. First, AMD product page updates moving from prototype language to orderable SKUs with PSU, dimensions, weight, noise, and OS support. Second, OEM announcements with two GPU versus four GPU chassis, RAM options from 512GB to 2TB, NVMe tiers, and onsite warranty terms. Third, ROCm release notes naming MI350P PCIe plus Halo topology with validated PyTorch and vLLM versions. Fourth, third party power plus thermal logs showing wall draw, inlet temps, and multi day stability, not just fit screenshots. Fifth, price leaks with BOM detail that confirm or break the 100K to 150K speculation band. Until three of those five land, treat Halo as directionally exciting but not actionable for purchase. If you must plan facilities now, reserve two dedicated circuits, a vented closet with 10GbE, and floor space for a 50kg plus tower. Those reservations cost little and prevent retrofit pain if you later order.
FAQ: 6 Answers
1. What is Threadripper Halo Station per VideoCardz Sept 4?
A liquid cooled deskside AI workstation shown at IFA 2026 with a 96 core Threadripper PRO CPU, two Instinct MI350P cards for 288GB HBM3E in the demo, path to four cards for 576GB, and up to 2TB DDR5 system memory. Goal is local trillion parameter class AI. Prototype status, no price or ship date confirmed.
2. How does workstation differ from desktop?
Workstation targets sustained multi day throughput, ECC paths, HBM residency, 24/7 liquid cooling, multi kW power, and ROCm Linux software. Desktop targets bursty gaming frames, 16 to 32GB VRAM, single circuit power, and game drivers. Halo is not a gaming value play. A gaming PC beats it per frame for one tenth the cost.
3. Can 576GB really hold a trillion parameter model?
At 4 bit quantization near 0.5 bytes per parameter plus overhead, weights need near 500GB, which fits in 576GB HBM3E with limited headroom for context. At 8 bit or 16 bit it does not fit without offload. Real deployments will use 4 bit quants, short to medium context, and careful batch sizing. The two GPU 288GB demo fits 200B to 500B class comfortably.
4. What power and room does it need?
Demo class likely draws 1800 to 2200W at the wall under load. Four GPU max likely exceeds 3000W. That exceeds one standard room circuit. Plan multiple dedicated circuits or lab feeds, vented closet or HVAC, 10GbE to desks, and floor rated for 50kg plus. Idle is far lower, but size for sustained load.
5. How much will it cost?
Unknown. AMD has not announced pricing. Outlets speculate 100,000 to 150,000 dollars based on CPU plus 2TB DDR5 plus accelerator street math. Treat as unconfirmed planning band. Require itemized OEM BOM with RAM, SSD, network, warranty, and field upgrade terms before budgeting to the dollar.
6. Should you buy, rent, or wait?
Buy path only for labs with private data, steady HBM need, Linux talent, and power plus cooling ready. Rent cloud HBM if needs are bursty or models are undecided. Wait if you need retail validation, OEM SKUs, and proven ROCm throughput. Most creators should buy 128GB unified mini or 24GB gaming plus cloud burst instead.
Written by Indie Kings, September 8, 2026. Source base: VideoCardz Sept 4, 2026 Halo Station report plus IFA 2026 keynote context. For live workstation plus GPU bundle prices see Bundle Tracker.
Related: Bundle Tracker | VideoCardz Sept 4 Halo Report | AMD Halo Station Page