// HACKER NEWS — CYBERSECURITY
Floppy Emu Hardware Failure Analysis Results
Nobody enjoys troubleshooting non-working hardware, and I’m no exception. Every time I ask a contract manufacturer to assemble a batch of more Floppy Emus, there are always a few that don’t pass QA testing. What happens with those? For several years the accumulating QA failures have been sitting in a pile in the corner of my office, along with a few customer returns, all waiting for the day when I would dedicate time to their investigation. It was a long time coming, but “Failure Analysis Day” finally arrived, or more like Failure Analysis Month, and the results were pretty interesting.
Here’s a breakdown of the unique causes of failure that I identified, and the number of boards affected by each cause.
This was a strange failure that required a lot of time to track down initially, but once I learned to recognize the symptoms, I realized that most of my QA failures were due to this problem. I wrote about the clock crystal mysteries in more detail in a separate post last month. The short story is that during my most recent manufacturing batch, my normal supplier of clock crystals was out of stock. The contract manufacturer, with my approval, substituted a different crystal with the same specs. It shouldn’t have affected anything. But somehow it did.
26 of the QA failures appeared to have no functioning clock at all. The board utterly failed to do anything when the microcontroller clock source was changed to the external crystal. A further 16 displayed some level of function, but with erratic behavior or failures at higher clock speeds. Initially I wasn’t sure whether this was due to the newer crystals exposing some defect or fragility in the Floppy Emu design, or whether it was simply caused by defective or damaged crystals. Eventually I came to the opinion that the crystals were damaged by rough handling or overheating during assembly. Replacing the crystals and reprogramming the boards resolved all the issues.
The Xilinx CPLD chip on the Floppy Emu is a delicate flower that has long been a source of challenges. Prior to this year’s crystal-gate debacle, the CPLD was the single biggest source of failures that I’d observed. As the chip that’s directly connected to the Floppy Emu’s external interface, it bares the brunt of any static discharge or electrical stress. It’s a 5V-tolerant 3.3V part, but its 5V tolerance has sometimes seemed a bit questionable, at least in the way it’s used here.
Symptoms of a failed or bad CPLD can include disk emulation failures, overheating, or erratic behavior. Usually the device will still be functional and text appears on its display, but the disk features no longer work. In extreme cases the failed CPLD acts as a hard short-circuit from power to ground, and then nothing works. Replacing and reprogramming the CPLD resolved all of the problems with these boards.
Amusingly, or depressingly depending on your perspective, the third leading cause of QA failures was that the microcontroller simply wasn’t programmed. Somebody fell asleep at the switch at the contract manufacturer, lost track of what they were doing, put a PCB in the wrong pile, or whatever. A board with an unprogrammed microcontroller will appear completely dead at first glance, but it still responds in the debugger and it only takes a few seconds to flash the chip and get everything working.
Every component must be electrically bonded to the PCB with solder. Soldering problems can be tough to spot with the naked eye, but usually jump out under magnification, so one of my first troubleshooting steps is usually to look at a problematic board at 10x. I really should get a cool desktop microscope, but for the moment I’m using a cheap 10x jeweler’s loupe which works well enough.
A couple of boards had too much solder in places, resulting in a solder bridge that unintentionally connected two adjacent IC pins. But it was more common to find joints with insufficient solder or poor solder joints, where an IC pin was sort of resting on