If you're planning a local AI inference build, the math changed underneath you this year. Per Tom's Hardware, average DDR5 kit prices are up 355-485% year-over-year, and some capacities are running roughly 10x their lowest-ever tracked price — 128GB of DDR5-6400 that once bottomed out at $329 now costs $3,399. This isn't a routine component-pricing blip; it's AI datacenter demand consuming the global memory supply, and it directly changes what a local-inference machine costs to build today.
TL;DR
| Question | Answer |
|---|---|
| DDR5 YoY price change? | +355% to +485% depending on speed/capacity (Aug 2025 → Aug 2026) |
| Worst-case multiple vs. historic low? | ~10x (128GB DDR5-6400: $329 low → $3,399 now) |
| Is DDR4 affected too? | Yes — up 120-177% YoY on spillover demand from priced-out DDR5 buyers |
| Cause? | AI datacenter/HBM demand locking up global DRAM fab capacity through 2027 |
| When does it ease? | Not expected before 2027-2028 at the earliest, per industry executives |
| What should local-AI builders do now? | Budget RAM as the dominant build cost, or default to cloud inference for now |
The actual numbers
Tom's Hardware's tracking, drawing on PCPartPicker historical pricing, shows the scale clearly. Comparing August 2025 to August 2026 average prices:
| Kit | Aug 2025 | Aug 2026 | Change |
|---|---|---|---|
| DDR5-4800 2x16GB | $90 | $425 | +372% |
| DDR5-5600 2x16GB | $116 | $528 | +355% |
| DDR5-6000 2x16GB | $108 | $572 | +429% |
| DDR5-5600 2x32GB | $191 | $1,118 | +485% |
| DDR5-6000 2x32GB | $222 | $1,272 | +473% |
DDR4 saw smaller but still steep increases (120-177% YoY across common capacities) as builders priced out of DDR5 shifted demand onto older platforms — meaning there's no cheap escape hatch to older hardware either. Best-vs-lowest-ever US pricing tells the same story from a different angle: a 96GB DDR5-6000 kit that bottomed at $189 now runs $1,799, and the 128GB DDR5-6400 kit mentioned above is at roughly 10x its historic floor. Germany's ComputerBase independently confirmed the same trend in Europe — average RAM prices there are up 345% since September 2025, with SSD and HDD prices both up over 125% in the same window.
Why: AI is buying the entire supply chain
The cause isn't mysterious. Hyperscale AI buyers have reportedly locked in nearly all global DRAM production capacity for 2027 through advance deposits, and DRAM's per-kilogram value is now, by some estimates, over half that of solid gold. All four major memory vendors — SK hynix, Samsung, Micron, and China's CXMT — have seen revenue double, triple, or more in the past year, because consumer and PC memory is now a lower-priority allocation compared to high-bandwidth memory (HBM), the memory type AI accelerators need, which competes for the same fabrication capacity as ordinary DDR5.
New fabs take years and billions of dollars to build. SK Hynix, Micron, Samsung, and Nanya all have new capacity under construction, but none of it is expected to meaningfully ease supply before roughly 2027-2028. SK Hynix's CEO has reportedly gone further, calling 2027 the worst year yet for memory supply industry-wide, with demand outstripping production into 2030.
HBM vs DDR5: why your laptop RAM and an H100 share a supply chain
The headline is about DDR5 kit prices, but the underlying squeeze is HBM — high-bandwidth memory stacked directly onto AI accelerator packages. HBM and DDR5 are not the same product, but they compete for the same scarce inputs: advanced DRAM fab lines, packaging capacity, and skilled yield engineering.
| Factor | HBM (AI accelerators) | DDR5 (PC / workstation) |
|---|---|---|
| Primary buyers | Hyperscalers pre-booking 2027 capacity | Consumers, small builders, enterprise desktops |
| Unit economics | Higher margin per wafer; long-term supply contracts | Lower margin; spot-market pricing |
| Fab priority | Vendors allocate leading-edge capacity here first | Gets whatever capacity remains after HBM orders |
| Packaging | Requires 3D stacking (TSV, CoWoS-class interposers) | Standard DIMM modules |
| Your build impact | Indirect — you never buy HBM, but you pay for what it displaced | Direct — every GB in your rig is priced against datacenter demand |
When SK hynix, Samsung, or Micron can sell a wafer's output as HBM to a hyperscaler under a multi-year contract, consumer DDR5 becomes the residual market. That is why DDR4 spillover rose 120-177% too: builders priced out of DDR5 did not find a cheap alternative; they just moved demand sideways onto an older standard that was never meant to absorb this volume.
For local-AI builders, the practical takeaway is that memory shortages will not self-correct when GPU prices fall. A cheaper RTX or M-series Mac does not help if the 128 GB unified memory configuration you need for a 70B-class model is still allocated away from the consumer channel.
Cloud vs local: a TCO table for August 2026
The RAM spike changes the break-even math. Below is a simplified comparison for a builder running a 70B-class open-weight model at moderate daily volume (~500K input tokens, ~100K output tokens per day). Cloud pricing uses typical API rates; hardware assumes a workstation build sized for that model class with current August 2026 component pricing.
| Line item | Cloud (API inference) | Local build (buy today) |
|---|---|---|
| Upfront hardware | $0 | ~$4,500 GPU + ~$2,800 CPU/motherboard/PSU/case |
| Memory (128 GB DDR5) | N/A (provider absorbs) | ~$3,399 (10x historic low on tracked kits) |
| Monthly inference cost | ~$450–650 at listed API rates | ~$35–50 electricity (24/7 idle + burst load) |
| 12-month total (Year 1) | ~$5,400–7,800 | ~$10,700–11,000 all-in |
| 24-month total | ~$10,800–15,600 | ~$11,500–12,200 (hardware + power) |
| Break-even vs cloud | — | Roughly 18–24 months at current RAM prices; was ~12 months in 2025 |
| Privacy / offline | Requires provider trust + network | Full control |
| Marginal cost per extra token | Linear with usage | Near-zero after hardware sunk |
Two nuances worth adding before you treat this as a spreadsheet decision:
- Apple Silicon changes the shape. Unified memory on an M-series Mac avoids the DDR5 DIMM market entirely, but you pay Apple’s configuration premium instead. DFlash 2 on MLX shows how speculative decoding can stretch what you get from fixed on-chip memory — useful if you already own the hardware, less useful if you are buying new at peak RAM pricing.
- Smaller models flip the math faster. If you are running 7B–14B models on 32–64 GB, cloud still wins on Year-1 TCO for most usage levels — see Kimi K3 on desktop hardware for what “right-sized” local inference looks like when memory is the binding constraint.
What people are asking
“Should I buy RAM now before it gets worse?” Only if you are actively blocked on a project today. Industry executives are signaling 2027 as the tightest year yet; buying at a 10x historic low is locking in the worst price of the cycle unless you have no alternative.
“Can I run big models without buying 128 GB?” Yes — quantization (Q4/Q5), disk-offloaded streaming, and speculative decoding all reduce resident memory. The trade-off is latency and throughput, not capability. Check whether your actual workflow needs full-precision context or just reliable answers.
“Does this affect cloud API prices?” Indirectly. Hyperscalers pre-booked capacity; consumer RAM spikes do not automatically raise token prices tomorrow. Longer term, if HBM and power stay expensive, frontier inference costs have less room to fall — which makes local builds look better on a longer horizon even at inflated upfront RAM cost.
“Is used RAM a workaround?” Second-hand DDR5 exists, but verify compatibility (EXPO/XMP profiles, QVL lists) and expect limited supply as enterprise refresh cycles slow. Used is not immune to the same demand curve — just discounted from an already-high floor.
What this actually costs a local-inference build
For anyone running open-weight models locally — via llama.cpp, Ollama, or similar — RAM has quietly become the dominant line item in the build, not a rounding error. A 64GB DDR5-6000 kit averaging roughly $190-220 last August now averages $1,100-1,270. If you were sizing a rig to comfortably run larger open-weight models (see explainx.ai's guides on running SOTA LLMs locally or Kimi K3 on desktop hardware), that memory alone may now cost more than the rest of a mid-range build combined.
Practical guidance if you're building anyway:
- Right-size to your actual model, not aspirational headroom. Check explainx.ai's top open-weight models for a laptop or efficient small models like GLM streaming on 25GB RAM before buying more memory than the model you're actually running needs.
- DDR4 platforms aren't a real discount right now given the 120-177% spillover increase — don't assume "older" means "cheap" in this market.
- Benchmark cloud/API cost against the inflated hardware price of ownership. For most builders who don't specifically need offline, private, or zero-marginal-cost inference, cloud is now the more rational default purely on cost — a calculation that's genuinely flipped from a year ago. explainx.ai's projection on local hardware trends through 2028 is worth reading if you're deciding whether to wait out the shortage.
- If you already own the hardware, don't upgrade RAM speculatively right now unless you're actively blocked — the price you'd pay today is close to the worst point in this cycle by every metric above.
Summary
DDR5 memory prices are up 355-485% year-over-year and as much as 10x historic lows, driven by AI datacenters locking up global DRAM production capacity through 2027. DDR4 offers no real escape either, up 120-177% on spillover demand. For local AI builders, this means memory now dominates build cost in a way it didn't a year ago, and no relief is expected before 2027-2028 at the earliest — a real, multi-year shift in the economics of running models on your own hardware versus in the cloud.
Builder takeaways
- Default to cloud for new projects unless offline, privacy, or zero marginal cost is a hard requirement — the RAM line item alone can add 12+ months to hardware payback.
- Right-size memory to the model, not the benchmark leaderboard. A 32 GB rig running a 14B Q4 model beats a $3,399 RAM bill for headroom you never use.
- Watch HBM allocation news, not just PCPartPicker charts — DDR5 relief tracks fab capacity freed from AI accelerator orders, not consumer demand falling.
- If you already own capable hardware, optimize before upgrading: DFlash 2 for Apple Silicon throughput, quantization guides for GPU builds, Kimi K3 locally as a reference stack for open-weight desktop setups.
Related on explainx.ai
- DFlash 2 on MLX: speculative decoding on Apple Silicon
- Kimi K3: run locally on open weights desktop hardware
- Top 10 open-weight models for a laptop
- Sentence Transformers v6: multi-vector ColBERT for RAG
- France sovereign AI: Mistral excludes OpenAI
- RadixArk Miles v0.1: open-weight RL post-training
- How to run open-source models locally with OpenCode
- Claude/Fable 5 local hardware projection through 2028
Source: Tom's Hardware — "Memory prices climb 500% in 12 months" (Zak Killian, August 2026), citing PCPartPicker historical pricing data and ComputerBase's European market tracking.
Pricing reflects data reported as of August 19, 2026. Component prices are volatile — verify current pricing before budgeting a build.
