ECC RAM: What It Is and Why Servers Use It

How error-correcting memory detects and fixes bit flips, why they happen more than you would think, what it costs, and how to check a server has it.

Published
Reading time
3 min

Memory is not perfect. A bit stored as 1 occasionally reads back as 0, because of a cosmic ray, a manufacturing defect, heat, or electrical noise from the module next door. On a laptop that is an unexplained crash once a year. On a server holding a database in memory for months, it is silent corruption. ECC memory is the fix, and it is one of the quiet differences between a server and a desktop.

What ECC does

Error-Correcting Code memory stores extra bits alongside each 64-bit word — typically 8 extra, making 72 — computed from the data. On every read, the controller recomputes the code and compares. A single flipped bit is detected and corrected on the fly; a double flip is detected and reported so the OS can act rather than continue with bad data.

Without ECC, a flipped bit is simply wrong data: a corrupted row in a database page, a mangled file in the page cache written back to disk, a pointer that sends a process into a crash.

How often bits flip

More than intuition suggests. Google's 2009 study of its own fleet found about 8% of DIMMs experiencing at least one correctable error per year, with error rates far higher on some modules than others. A server with 128 GB of RAM running for a year can expect several corrected errors. Each of those, on non-ECC memory, would have been a corruption you never saw.

Rowhammer — deliberately flipping bits by hammering adjacent rows — turned this from a reliability topic into a security one; ECC raises the bar for those attacks substantially.

What it costs

  • Price: ECC modules cost 10–20% more, and the platform (CPU and board) must support them, which is where most of the cost difference between server and desktop parts comes from.
  • Speed: a few percent, from the extra check on each access. Not measurable in any web workload.
  • Capacity: none visible; the extra bits are on the module.

Registered vs unbuffered

Server ECC memory is usually RDIMM (registered): a buffer chip on the module lightens the electrical load on the memory controller, allowing more modules per channel and so more total RAM. Desktop and small-server ECC is UDIMM (unbuffered). They are not interchangeable; the board decides.

Which CPUs support it

All AMD EPYC and Intel Xeon. AMD Ryzen supports ECC on many boards (the "Pro" variants officially). Intel's consumer Core chips do not, except some entry Xeon-E parts. So a "server" built from desktop parts — which some budget hosts do — is likely running without ECC. Ask.

Check whether a server has it

On Linux:

bash
sudo dmidecode -t memory | grep -E "Error Correction Type|Type:|Size"

Error Correction Type: Multi-bit ECC is what you want; None means non-ECC. And to see whether errors are being caught:

bash
sudo apt install rasdaemon -y && sudo systemctl enable --now rasdaemon
sudo ras-mc-ctl --error-count

A module that reports a rising count of corrected errors is one to replace before it produces an uncorrectable one. That is the other benefit of ECC: it tells you which module is going bad while everything still works.

On a VPS you cannot see the host's memory — dmidecode shows the virtual machine's fake DIMMs. What you can do is ask the provider. VPSPioneer's managed VPS hosts run registered ECC memory on EPYC platforms, and corrected-error counts are one of the hardware metrics we watch.

Where it fits among the other parts

ECC memory, mirrored NVMe storage and dual power supplies each remove one category of silent failure. The rest of the server's insides, and which ones change what you experience as a customer, are in inside a server.

#hardware#ram#ecc#reliability#servers

Keep reading

More from Hardware & Datacenter

All guides

Hardware & Datacenter

RAID Levels Explained: 0, 1, 5, 6 and 10 for Servers

What each RAID level does with your drives, how much capacity and failure tolerance you get, why hosting servers use RAID 10, and why RAID is not a backup.

3 min read →

Hardware & Datacenter

Inside a Server: CPU, RAM, Storage and NIC Explained

What each component in a hosting server does, how server parts differ from a desktop, and which spec-sheet numbers actually change site performance.

3 min read →

Hardware & Datacenter

NVMe vs SATA SSD vs HDD: What the Difference Means for Hosting

How the three storage types differ in latency and IOPS, why a website's workload favours NVMe by a wide margin, and when the difference is not worth paying for.

3 min read →