Hardware Troubleshooting
Hardware diagnosis is controlled uncertainty reduction. The fastest troubleshooters do not guess more quickly; they create a reliable symptom, divide the system into smaller possibilities, and change one variable at a time.
TL;DR
- Record the exact symptom and conditions before changing components.
- Protect data and follow the device's service instructions.
- Use compatible known-good parts to test a specific hypothesis.
- Repeat the failing workload and adjacent checks after repair.
Quick Example
This example shows a diagnostic record, not a universal instruction to stress-test hardware.
Core Concepts
POST Versus Operating-System Failure
Power-on self-test (POST) happens before the operating system loads. Firmware lights or beep codes must be interpreted using that model's manual. If firmware works but the OS does not boot, investigate the boot device, configuration, and software as well as hardware.
Signals Are Evidence
An error counter, temperature reading, or memory-test failure narrows a diagnosis. Interpret it with workload and time: an old logged error need not explain today's symptom.
Start With the Symptom
Record exactly what happens, when it began, what changed, whether it is repeatable, and which states trigger it: cold boot, sustained load, idle, sleep, movement, battery, or a particular peripheral. Photograph error lights and capture firmware or operating-system messages before power cycling.
Divide the Problem
Reduce to a minimum configuration, reseat only after inspection, swap with a compatible known-good component, and test the suspect component elsewhere when safe. Follow the manufacturer's shutdown, power-disconnection, battery, and electrostatic-discharge procedures before internal work. Do not open power supplies or handle damaged batteries; refer those repairs to qualified service personnel. A replacement that appears to work is evidence, not proof; the act of moving cables or cooling the system may have changed the condition.
Protect data before running stress tests. A failing drive with irreplaceable data needs a recovery decision before repeated scans, repair writes, or power cycles. SMART warnings are useful evidence, but a passing health report does not prove a drive is healthy.
Close the Repair Properly
Run the workload that exposed the fault, then test adjacent functions. Check temperatures, error counters, storage health, memory diagnostics, firmware logs, sleep and wake, and sustained operation. Record the failed part, observed evidence, repair, validation, asset update, and any data or warranty action.
Best Practices
Establish a Baseline
Record cable positions, firmware settings, component identifiers, and the initial result. This makes the previous state recoverable and prevents a helpful change from hiding the cause.
Define Stop Conditions
Stop testing when there is smoke, liquid damage, a swollen battery, unsafe heat, or worsening storage symptoms. Repeated power cycles are not a recovery strategy.
Common Mistakes
Treating a Health Check as Proof
Bad: Declare storage healthy because SMART passes.
Correct: Correlate logs, latency, device behavior, and data-integrity symptoms.
Swapping Several Parts at Once
Bad: Replace memory, storage, and cables together and claim the RAM was faulty.
Correct: Change one safe variable, record the result, and reproduce the original failure condition.
FAQ
Can software look like a hardware fault?
Yes. Drivers, power-management settings, firmware, and resource pressure can cause crashes or disconnects. Compare safe configurations before concluding a part has failed.
Should I run diagnostics on a failing drive?
Protect recoverable data first. If the data is irreplaceable or symptoms are severe, seek a recovery assessment before scans or write-based repair.
When is a repair complete?
When the original symptom no longer occurs under representative conditions and adjacent functions still work. Record the evidence and any remaining uncertainty.