DIMM (Memory) Error Troubleshooting¶
Purpose¶
This document explains how to identify, confirm, and resolve DIMM (memory) errors using the iDRAC console, system logs, POST messages, and physical troubleshooting.
Prerequisites¶
- iDRAC access
- Console access (Virtual Console or Physical KVM)
- Proper ESD precautions
- Server powered safely
Step 1: Identify the Issue in iDRAC Console¶
- Log in to the iDRAC console.
-
Navigate to:
-
System Overview
- Hardware / Memory Section
-
Check system health status:
-
Look for Memory / DIMM warnings or critical alerts.
-
Note:
-
DIMM slot number (example: A1, B2, C1)
- Error description
- Timestamp of the error
Purpose: To confirm whether the issue is related to memory and identify the affected DIMM slot.
Step 2: Collect Logs from iDRAC¶
-
From iDRAC, collect the following logs:
-
Lifecycle Logs
- System Event Logs (SEL)
- Save the logs for reference.
-
Verify:
-
Any DIMM-related error entries
- Repeated or persistent memory errors
Purpose: Logs help confirm whether the error is intermittent or persistent and are required for escalation.
Step 3: Perform Power Drain¶
- Gracefully power off the server.
- Remove all power cables from the server.
- Presee the power on button Wait for 10 to 30 seconds (minimum).
- Reconnect the power cables.
- Power on the server.
Purpose: Power drain clears residual power and temporary hardware faults.
Step 4: Observe POST Logs During Boot¶
- Open the Virtual Console or connect to physical console.
- Power on the server.
-
Carefully observe the POST screen:
-
Memory initialization messages
- Any DIMM or memory-related errors
-
If:
-
Server fails to boot OR
- Any error appears during POST Then:
- Allow the server to complete power-up
- Wait 2 additional minutes
- Recheck logs in iDRAC
Purpose: POST messages give real-time confirmation of hardware issues.
Step 5: Confirm DIMM Error¶
-
Compare:
-
POST error messages
- iDRAC alerts
- System logs
- Confirm that the issue is specifically a DIMM error.
Purpose: Avoid unnecessary hardware replacement by confirming the root cause.
Step 6: Power Off and Open the Server¶
- Power off the server.
- Remove all power cables.
- Follow ESD safety guidelines.
- Open the system cover.
Purpose: Prepare for physical inspection of memory modules.
Step 7: Verify and Reseat DIMMs¶
- Locate the DIMM slot showing the error (example: A1).
-
Check:
-
DIMM is fully inserted
- Locking clips are properly engaged
-
Reseat the DIMM:
-
Remove the DIMM
- Reinsert it firmly into the same slot
-
Verify all DIMMs:
-
Properly seated
- Correct population order as per system guidelines
Purpose: Loose or improperly seated DIMMs are the most common cause of memory errors.
Step 8: Power On and Observe Again¶
- Close the system cover.
- Reconnect power cables.
- Power on the server.
-
Observe:
-
POST screen
- Any memory error messages
- Allow the server to boot fully.
- Check iDRAC logs and console.
Purpose: Confirm whether reseating resolved the issue.
Step 9: Swap DIMM Positions (Isolation Test)¶
If the DIMM error still appears:
- Power off the server.
- Remove all power cables.
- Open the system.
-
Move the DIMM to another slot:
-
Example:
- Original error: A1
- Move DIMM from A1 → B1
- Close the system cover.
- Power on the server.
- Observe:
-
POST screen
- Console messages
- Wait until the server fully powers up.
- Check iDRAC logs again.
Purpose: To determine whether the issue is with the DIMM slot / motherboard.
Step 10: Analyze the Result¶
Case 1: Same DIMM Slot Shows Error Again¶
- Error remains on the same slot (example: A1)
- DIMM moved, but slot still reports error
✅ Conclusion:
- Slot or motherboard issue ➡️ Escalate to support / vendor team
Case 2: Error Moves with the DIMM¶
- Error now appears in new slot (example: B1)
- Error follows the DIMM
✅ Conclusion:
- DIMM is faulty ➡️ Proceed with DIMM replacement
Final Summary (Quick View)¶
- Identify error in iDRAC
- Collect logs
- Power drain (2 minutes)
- Observe POST logs
- Reseat DIMM
- Swap DIMM positions
-
Decide:
-
Slot issue → Support team
- DIMM issue → Replace DIMM