Skip to content

DIMM (Memory) Error Troubleshooting


Purpose

This document explains how to identify, confirm, and resolve DIMM (memory) errors using the iDRAC console, system logs, POST messages, and physical troubleshooting.


Prerequisites

  • iDRAC access
  • Console access (Virtual Console or Physical KVM)
  • Proper ESD precautions
  • Server powered safely

Step 1: Identify the Issue in iDRAC Console

  1. Log in to the iDRAC console.
  2. Navigate to:

  3. System Overview

  4. Hardware / Memory Section
  5. Check system health status:

  6. Look for Memory / DIMM warnings or critical alerts.

  7. Note:

  8. DIMM slot number (example: A1, B2, C1)

  9. Error description
  10. Timestamp of the error

Purpose: To confirm whether the issue is related to memory and identify the affected DIMM slot.


Step 2: Collect Logs from iDRAC

  1. From iDRAC, collect the following logs:

  2. Lifecycle Logs

  3. System Event Logs (SEL)
  4. Save the logs for reference.
  5. Verify:

  6. Any DIMM-related error entries

  7. Repeated or persistent memory errors

Purpose: Logs help confirm whether the error is intermittent or persistent and are required for escalation.


Step 3: Perform Power Drain

  1. Gracefully power off the server.
  2. Remove all power cables from the server.
  3. Presee the power on button Wait for 10 to 30 seconds (minimum).
  4. Reconnect the power cables.
  5. Power on the server.

Purpose: Power drain clears residual power and temporary hardware faults.


Step 4: Observe POST Logs During Boot

  1. Open the Virtual Console or connect to physical console.
  2. Power on the server.
  3. Carefully observe the POST screen:

  4. Memory initialization messages

  5. Any DIMM or memory-related errors
  6. If:

  7. Server fails to boot OR

  8. Any error appears during POST Then:
  9. Allow the server to complete power-up
  10. Wait 2 additional minutes
  11. Recheck logs in iDRAC

Purpose: POST messages give real-time confirmation of hardware issues.


Step 5: Confirm DIMM Error

  1. Compare:

  2. POST error messages

  3. iDRAC alerts
  4. System logs
  5. Confirm that the issue is specifically a DIMM error.

Purpose: Avoid unnecessary hardware replacement by confirming the root cause.


Step 6: Power Off and Open the Server

  1. Power off the server.
  2. Remove all power cables.
  3. Follow ESD safety guidelines.
  4. Open the system cover.

Purpose: Prepare for physical inspection of memory modules.


Step 7: Verify and Reseat DIMMs

  1. Locate the DIMM slot showing the error (example: A1).
  2. Check:

  3. DIMM is fully inserted

  4. Locking clips are properly engaged
  5. Reseat the DIMM:

  6. Remove the DIMM

  7. Reinsert it firmly into the same slot
  8. Verify all DIMMs:

  9. Properly seated

  10. Correct population order as per system guidelines

Purpose: Loose or improperly seated DIMMs are the most common cause of memory errors.


Step 8: Power On and Observe Again

  1. Close the system cover.
  2. Reconnect power cables.
  3. Power on the server.
  4. Observe:

  5. POST screen

  6. Any memory error messages
  7. Allow the server to boot fully.
  8. Check iDRAC logs and console.

Purpose: Confirm whether reseating resolved the issue.


Step 9: Swap DIMM Positions (Isolation Test)

If the DIMM error still appears:

  1. Power off the server.
  2. Remove all power cables.
  3. Open the system.
  4. Move the DIMM to another slot:

  5. Example:

    • Original error: A1
    • Move DIMM from A1 → B1
    • Close the system cover.
    • Power on the server.
    • Observe:
  6. POST screen

  7. Console messages
  8. Wait until the server fully powers up.
  9. Check iDRAC logs again.

Purpose: To determine whether the issue is with the DIMM slot / motherboard.


Step 10: Analyze the Result

Case 1: Same DIMM Slot Shows Error Again

  • Error remains on the same slot (example: A1)
  • DIMM moved, but slot still reports error

✅ Conclusion:

  • Slot or motherboard issue ➡️ Escalate to support / vendor team

Case 2: Error Moves with the DIMM

  • Error now appears in new slot (example: B1)
  • Error follows the DIMM

✅ Conclusion:

  • DIMM is faulty ➡️ Proceed with DIMM replacement

Final Summary (Quick View)

  • Identify error in iDRAC
  • Collect logs
  • Power drain (2 minutes)
  • Observe POST logs
  • Reseat DIMM
  • Swap DIMM positions
  • Decide:

  • Slot issue → Support team

  • DIMM issue → Replace DIMM