SGI Indigo R4400 running IRIX 6.5 with its matching retro display, three hours into the first validation run after the power supply repair

My Indigo R4400 is the prize of the collection, and it died a few months back. I had started it to do some work and it booted up fine. I left for about an hour, and when I came back it was dead. A power cycle did nothing except a single click from the speaker, and the LED on the main board never lit, which pointed to at least the 5V rail not being energized.

I had faced this problem before with my Indigo R3000. I did a lot of legwork trying to fix that supply, but in the end there just wasn’t much information on the internet about it, and certainly no schematics. That machine ended up with a new supply constructed from three different donor supplies. It worked, but it was a hack that consumed nearly the entire power supply location plus the drive bay. Not elegant.

The supply is an ITT PowerSystems PEC4044B (assembly 6202805), and it is basically unobtainable. None of the usual SGI parts sources turn one up, and I’ve been watching eBay for about four years to no avail.

There is a commercial supply you can put in the machine, a unit from a kind of medical device: the Power One PFC375-4000F industrial AC-DC supply. It has been used on older SGI gear, but they are very expensive and just another fill-the-drive-bay solution.

I decided that would be my ultimate fallback, and this time I would try a recap first. Also different this time: I used Claude, the AI tool from Anthropic, to help, which really turned out to be the X factor in fixing this supply. I started by feeding it everything I could find online (not much) plus my previous work and the things I knew about how the supply behaves from constructing the franken-supply, which carried over more than I expected. This is the write-up. Spoiler: the power stage was never broken. The supply was being shut down on purpose, by its own protection circuitry.

The PEC4044B is essentially undocumented. What exists online is a handful of forum threads from a few determined owners, JeffC’s recap thread on IRIXnet being the best of them, and none go very deep on how the thing actually works (I’m not sure anyone really knows). It is also remarkably small for a supply that delivers something like 35 amps at +5V, which is exactly what makes it nearly impossible to substitute: nothing modern in this form factor comes close on the 5V rail. Inside, it is two boards: the main power board, and a control board that plugs into an edge connector at a right angle and is further tied to the main board by two heavy wire bundles. That is about the extent of the documented anatomy; everything else in this article had to be worked out on the bench.

How it failed

This wasn’t a sudden death. The machine ran fine one day, started normally the next, then died unattended about an hour in, warm, and clicked once ever after. That detail (dies warm, after running a while) turned out to be the whole story.

There was even a warning shot. The R4400 write-up on this site records me noticing faint noise on the screen, suspecting a tired capacitor in the supply, and musing about a proactive recap. I never got around to it. I did search from time to time for a power supply repair company, and sent emails, but never got replies.

Fair warning

Some of this is mains-side work. A switching supply stores lethal charge in its bulk capacitors, and unplugged does not mean safe: the big pair of capacitors in this supply sits at around 300V combined and can hold it long after the cord comes out. In practice I found that mine generally discharged on their own pretty quickly, so there is clearly something in the design to drain them, but don’t assume that circuit is working.

The ritual before touching anything, every time: unplug, discharge the bulks through a 10kΩ 5W resistor on insulated clip leads, then verify with a meter that they’re down to a few volts.

If you use a scope (which turned out to be the key instrument in this diagnosis), keep it on the isolated secondary side; a grounded probe clip on the primary is a short circuit through your bench. If any of this sounds like news, do the reading on mains safety before opening one of these. The heat lamp described below protects the supply from mistakes. It does not protect you.

Tooling up

Some honesty about equipment first. I have a very good soldering station, but it is not built for high-power work, and this board has heavy pours and big through-hole parts. So I bought a Weller W100PG, the classic 100 watt spade-tip iron with a CT6F7 700°F tip, apparently a fixture on power supply benches for decades. For the post-mortem on whatever came off the board I added a Peak Atlas ESR70, a highly rated capacitance and ESR meter (the purple one that looks like a game controller) that discharges the cap automatically before measuring it. My advice, though: do the discharge yourself with the clip-lead resistor tool rather than subjecting the meter to a potentially large charge.

I also own a proper load bank and decided not to use it. The main reason: +5 is carried off this board by eight separate wires, so wiring a load bank in would have required a pretty elaborate connector scheme. It turned out to be a good call anyway, because the incandescent load bulbs told part of the story all by themselves, flickering visibly whenever the supply’s regulation got unsteady, a detail an electronic load would have swallowed silently.

The test jig

I bench-ran the supply outside the machine, because you really have no other choice.

Having messed with one of these before, I knew how difficult this supply is to work on; everything about it is compact. So in the workshop I built a wood jig that mounts the main supply board and lets me turn the whole assembly around to get at its different sides. One of the harder things about working on a PSU is that you kind of need three hands, and the jig takes the place of one of them.

While I was at it I added three light bulb sockets to the base: one that the incoming AC runs through, one for +5V, and one for +12V. Supplies generally need some kind of load on them to function correctly, and some will shut themselves down without one, feigning their own death.

The jig’s electrical parts were as follows:

  • A series ballast on the AC input. I started with a 100W incandescent dim-bulb and later switched to a 250W heat lamp when it turned out the 100W bulb starves this supply (more below).

  • A dummy load of automotive 1156 bulbs on +5 and +12, wired into the output connector with 18AWG solid core. That gauge turns out to be the magic size for this big connector: it seats snugly in the pins (20AWG is too loose), and short lengths of it made good test points for jumpering and probing signals like Soft Start and the pin 12 status output.

  • The Soft Start pin jumpered to ground to start the supply.

  • A Tektronix TBS1052B scope on the outputs. Not technically required, but it unlocked what was actually happening; without it, this repair would have been a guess.

Most of this is the standard kit from Bench Tools for Vintage Computer Repair; the 1156 bulbs and the heat lamp were improvised for this job.

Plywood bench jig holding the PEC4044B main board upright, with 1156 bulb dummy loads and the 250W heat lamp ballast mounted alongside
The jig: the main board held upright in plywood, 1156 bulb loads on +5 and +12, and the 250W heat lamp ballast

One wrinkle worth recording: the jig ended up holding the output connector upside down. I had a pinout from the R3000 franken-supply work, drawn looking into the chassis side, but on the bench I was facing the opposite connector, flipped upside down. Rather than re-derive every probe placement in my head and eventually get one wrong, I drew the second view and kept it on the bench. Both views are below; pin 11 is Soft Start, and pin 12 is the status output that later cracked the case.

Pin map of the PEC4044B 24-pin output connector in two views: looking into the chassis side, and looking into the cable side upside down as mounted on the jig
The output connector pin map: as documented during the R3000 work (left, looking into the chassis side) and as actually faced on the jig (right, cable side and upside down)

Why a heat lamp

The heat lamp is a dim-bulb tester scaled up to match the supply. A series incandescent ballast limits fault current: a hard short on the supply primary lights the bulb up bright and keeps it there, turning a mistake into light instead of smoke. But the ballast has to be sized to the load. A standard 100W bulb in series with a roughly 300W supply drops so much voltage that the supply browns out, and my early tests produced confusing dynamics (bulk caps starving, rails coasting on stored energy) that were artifacts of the bulb, not the fault. A 250W infrared heat lamp is just a big incandescent filament, with the same protective behavior but enough headroom that the supply saw full line (150V and 156V across the bulk caps, verified) while keeping the short-circuit protection. In practice, with a healthy supply, the infrared bulb doesn’t really even light up, though you can feel some heat coming off it.

Those numbers deserve a word, because they puzzled me at first. The bulk caps are the two big primary-side reservoirs that store rectified line voltage, and on this supply the pair (two 200V parts) sits in series as a voltage-doubler input, a common trick for running a big supply from a 120V outlet. So roughly 300V across the stack, a bit over 150V on each cap, is exactly what full line should look like. It’s also why the safety warning above is in here, and one reason for the wood jig: touch the wrong thing or create an accidental short and it is going to be a show.

The lamps were also free diagnostics in themselves. On the ballast side, a brief flash at plug-in and then back to near-dark is normal inrush charging the bulks, no short. On the output side, the brake-light bulbs glow steadily in proportion to load: the supply is running and regulating. And at the moment of each mysterious shutdown, the +5 and +12 bulbs went dark while the ballast stayed dark too. The supply had stopped drawing entirely, which meant the collapse was not a short or an overload dragging things down but the converter being commanded to stop. That one observation, before the scope ever confirmed it on pin 12, was the first evidence that the supply was entering standby rather than dying.

The test rig shopping list

The complete rig, as actually used, sourced almost entirely outside of electronics suppliers.

Hardware store

  • One 250W infrared heat lamp bulb: the series ballast, for the reasons above.
  • One keyless lamp holder rated for the heat lamp’s wattage, about $4 with two screw terminals. Porcelain, not plastic, at 250W.
  • One cheap three-prong extension cord you are willing to cut.
  • Wire nuts or crimp connectors, electrical tape, and heat-shrink tubing to keep everything tight, together, and insulated.

Auto parts store

  • 12V type 1156 bulbs (single-filament brake/backup bulbs) with pigtail sockets: the dummy loads for +5 and +12.

Electronics

  • The replacement caps themselves (the full order is below).
  • One 10kΩ 5W resistor on insulated clip leads: the cap discharge tool from the safety ritual above. Not optional.
  • Alligator clip jumper leads.
  • 18AWG solid-core wire for the output connector test leads.

Chasing it

First runs on the 100W bulb: the supply starts, regulates 5.0V for a few seconds, dies. No hard short (the ballast never flares bright; the +5 and +12 bulbs simply go dark). An immediate restart fails; wait a minute or two and it starts again, then repeats the behavior.

Things that got ruled out, in order:

Input side. On the 250W heat lamp the bulk caps measured 150V and 156V during the run, so full line was being delivered. It still died.

Sagging rails, undervoltage, light-load overvoltage. The scope showed +5 and +12 flat-topped and in regulation right up to a clean, simultaneous, abrupt collapse. No sag, no drift, no overload signature. My digital multimeter had earlier shown “9.x volts” on the 12V rail, which sent me chasing a sagging rail that didn’t exist. That was an artifact: a cheap DMM updates about every 1.2 seconds, and pointing it at an event that lives for under two seconds means it catches the rail mid-collapse and prints a voltage that never existed. Moral: get a scope. Lesson one about instruments lying.

Minimum-load shutdown. As mentioned above, some switching supplies deliberately shut down if they see too little load, and the bulb loads are small compared to what an Indigo draws. Ruled out because the trip wasn’t deterministic: the exact same bulb load that died in seconds would later run indefinitely. A real load threshold trips every time, not sometimes.

Then I went on a bit of a wild goose chase, and it matters because it produced a false cure. In the machine, Soft Start is grounded to tell the supply to run; on the jig I had been doing that with a jumper. Surmising that a marginal connection or a floating Soft Start might be the problem, I hard-jumpered it to ground, and the supply started working. Completely. It stayed on, survived power cycles, and ran for days. But it kept working even after the jumper came off, so the jumper had explained nothing and fixed nothing. A marginal component somewhere had simply been perturbed into behaving for a while. I didn’t trust it, and I kept the jig set up.

Reproduction and the smoking gun

A week later, a burn-in test brought the fault back: under 10 minutes on the bench and the supply would shut down. A wait of about five to ten minutes would let it start again. So it really hadn’t fixed itself; the problem had just moved its timing.

Two facts about the shutdown itself mattered. The heat lamp stayed dark at shutdown, meaning the supply stopped drawing power; this was not a short. And an immediate restart always failed while a one-to-two minute unplug always recovered it. That’s a latch holding state on stored charge in the housekeeping supply (the small always-on section, itself riding on a capacitor, that powers the control logic even when the main outputs are off), not a component cooling off.

The control connector has a status output (pin 12 in my pin map, low = standby). I put one scope channel on +5, one on pin 12, set a single-shot trigger on the falling rail, and let it soak.

Tektronix TBS1052B single-shot capture: the blue pin 12 status line snaps low in one edge while the yellow +5 rail decays as an RC coast
The moment it confessed. Blue: pin 12 status snapping low. Yellow: +5 coasting down as the output caps drain into the bulb

Pin 12 snaps low in one crisp edge as the collapse begins, and +5 decays as a passive RC coast, just the output caps draining into the bulb. The converter stopped switching. It wasn’t dragged down; it was commanded into standby. And the pull-up on pin 12 is powered by the housekeeping supply on the control board that demonstrably stays alive after shutdown (it’s what holds the lockout latch), so that low is driven, not an artifact. It won’t hold forever; it’s being fed from a cap on the control board, which is exactly why a couple of unplugged minutes clears the lockout.

That’s the whole fault: a thermally-drifting component in the control board’s standby/supervision path falsely tripping the run/standby latch. On this board that path runs through a bank of LM339s, garden-variety comparator chips whose whole job is watching voltages and voting on whether the supply should be allowed to run. The protection circuit was the killer. Everything else, the clean collapses, the lockout folklore, the warm-death timer, the false healing, falls out of that one mechanism.

A key factor in all of this is that Claude figured most of it out. When I asked how it knew what those components do, on a supply with no schematics in existence, the answer was that this is a pretty classic power supply design from the 90s. It had simply inferred the architecture from the parts and the behavior.

It even explains the single click at power-on, though this part is a guess: the Indigo’s speaker sits behind an amplifier running on 12V, and the click would be the pop of the rails arriving and being commanded away again a heartbeat later. The machine wasn’t refusing to start. It was starting, and being shut down.

The fix

User JeffC on IRIXnet had recapped the same model and documented the capacitor list in thread 4686, which user legodude turned into a clickable DigiKey cart. I ordered it. One warning: verify values against your own board. JeffC’s board (part 6064459) has C127 = 47µF 50V; mine (6202805) is 120µF 35V. Same model number, different build.

I decided to recap the control board first, since that’s where the supervision circuitry lives, and because Claude had been convinced almost from the first scope traces that the control board was the culprit. There was a discipline to it: recap one board, reassemble, prove it against the untouched main board, and only then move on. You want to control the blast radius.

What I did differently this time: before touching anything I made meticulous diagrams. Pinouts, cap locations, cap orientations (the blue marks on the cap map below show which way each negative terminal faces). In spite of all that, one cap still went in backwards and one went into the wrong location. Both were caught before power-up, one by inspection and one by in-circuit measurement with the ESR70. Diagrams don’t prevent mistakes; they make them findable.

The jig made retesting cheap enough that you could honestly do it between every single replacement. In practice I batched nearly the whole control board in one pass: the 22µF trio C232, C253, and C259 (housekeeping/bias, and the prime suspects; these had gooey leads, meaning electrolyte had seeped past the bung, the cap’s rubber end seal), plus C236 (47µF) and C235 (4700µF). C233 stayed original; its replacement was still on backorder. I scrubbed every site with isopropyl alcohol. Leaked electrolyte is mildly conductive, and a residue film near high-impedance comparator inputs is itself a candidate trigger.

Result: after 60 minutes of run time on the bench, the supply was still going, and it restarted cleanly after that, warm, the exact maneuver that used to poison the latch. A byproduct I noticed: in tests before the control board recap, the light bulbs would shift slightly in brightness from time to time. I believe that was the regulation hunting, one more indication that something on the control board was not working correctly. After the recap, rock steady.

So the entire supply worked after touching only the control board. I couldn’t test every cap I removed (several were destroyed coming out, and I had no meter that could do in-circuit ESR tests beforehand), but the evidence that something on that board was responsible is pretty good. Everything after this point is preventive maintenance.

Which cap was it?

I can’t tell you. Every cap that survived extraction intact measured healthy on the ESR70. C235 read 4894µF and 0.05Ω ESR at 33 years old, factory-fresh. The prime suspects, the 22µF trio, got mangled coming out and couldn’t be measured. So: circuit convicted, no individual confession. Caps age individually. Measure what you can, replace the implicated neighborhood, and prove the cure with a soak test rather than a story.

The main board (preventive)

While the jig was set up: C105 (2.2µF 400V, primary side), C106 (originally 47µF 25V, the supply reservoir for the UC2845, the PWM controller chip that actually runs the converter; I put the best cap in the order here, a Panasonic FR 47µF 50V at 0.27Ω ESR), and the two 200V 1200µF bulks. Everything on the main board was actually working; this was me wanting to do them all while the supply was open. The order of operations was partly forced, too: C105 and C106 sit where you can’t really work on them until the big bulk cans are out of the way, so C108 and C109 came out first.

The big bulks are worth documenting: their designators are C108 and C109, hidden under the cans and, as far as I can tell, never recorded anywhere online until now.

A warning about extraction on this board: the wide heat-sinking traces drink heat, and the old joints will not let go for ordinary desoldering. The only sequence that worked was to flood the joint with fresh solder, wick it off, flood it again, wick again, and finally walk heat back and forth between the leads while pulling gently on the can from the other side of the board. The jig earned its keep here, holding the main board upright with a free hand on each side. And in spite of all that, I still ripped off a trace: 2.5mm wide off C109, top side, running from the cap’s through-hole to a neighboring resistor R104. A few pads came off too.

The C109 site on the main board with the broken trace running from the through-hole to a neighboring resistor
The C109 crime scene: the broken trace runs from the through-hole to the resistor. The bodge wire runs lead-to-lead on the back side

The saving grace of this board is also its trap. It is two-layer, component side and back side, and usually both traces for a given joint run on one side or the other. The holes are plated through, so when a destroyed pad sits on the side with no trace, nothing is lost; the barrel still carries the connection, and every such site here continuity-verified fine. But when the pad or trace on the connected side goes, you are down to having to install a bodge wire. For the C109 break I ran insulated 20AWG from the cap terminal on the back side of the board lead-to-lead to the resistor leg. High power boards are generally built this way.

Closeup of the green insulated bodge wire emerging beside the C109 silkscreen and soldered to the resistor leg
The bodge up close: through from the back side at C109, soldered to the resistor leg where the broken trace used to land

The two big caps originally had a plastic-like shroud around their sides, and it got destroyed during removal. You want to replace it: when the control board is installed it sits very close to these two caps, and they put that insulator there for a reason. I ended up using Pangda adhesive insulating paper (0.2mm fish paper, in green), formed around the new cans in two layers.

New C108 and C109 bulk capacitors wrapped in a green adhesive fish paper barrier on the main board, with the green bodge wire landing on a resistor leg nearby
The new bulks in their rebuilt green barrier. The green bodge wire arriving from the far side of the board and landing on the resistor leg is the trace repair

What didn’t get replaced

I wanted to do them all, but pulling C108 and C109 is what broke the traces, and after patching that damage I got cold feet. The board already worked after the control board fix, and every extraction on a 33-year-old board is optional surgery with real risk. So the last caps were saved for another day: the four 3900µF 10V output filters (C118 through C121), C127, and C233.

The output filters, at least, I could clear on the evidence from the scope. Extraction on their heavy 5V pour joints is where this board’s pads die, there was zero evidence against them, and output filter health is directly measurable: AC-coupled scope on the +5 output.

Two traps in that measurement are worth passing on. At slow sweep speeds in scan mode the display draws a min/max envelope; the rail looks like 400mV of grass and it’s meaningless. And most of what you see even at proper settings (10µs/div, bandwidth limit on) is the probe’s ground-clip loop picking switching edges out of the air. The control experiment: short the probe tip to its own ground clip. If the picture doesn’t change, you’re looking at antenna pickup, not ripple. Mine didn’t change. The real ripple was below the noise floor, a flat baseline between switching-edge bursts even at 200mV/div. The bank is fine. It stays, verified, and the new caps go in the drawer as spares for the next supply.

Tektronix TBS1052B showing the AC-coupled +5 output at 200mV per division: sharp bursts at the switching edges with a flat baseline between them
The +5 output, AC-coupled at 200mV/div. The sharp bursts line up with the converter's switching edges, and per the null test they are mostly ground-clip pickup. The flat baseline between them is the actual rail: tightly regulated, ripple below the noise floor. This is what a healthy output filter bank looks like

The cap map

Every electrolytic on both boards, with original values. Verify every value against your own silkscreen before ordering; this is assembly 6202805, and JeffC’s 6064459 differs in at least one value.

Annotated photo of the PEC4044B main board and control board with every electrolytic capacitor labeled with its designator and value, the replaced ones checked, and blue polarity marks showing negative terminal orientation
The cap map: both boards, every electrolytic. The green check marks signify the caps I ended up changing, and the blue marks record which way each negative terminal faces
DesignatorBoardOriginal valueReplacedNotes
C232, C253, C259Control22µF 25VYesHousekeeping/bias; the prime suspects, gooey leads
C235Control4700µF 16VYesOld part measured 4894µF, 0.05Ω ESR
C236Control47µF 25VYes
C233Control1000µF 50VNoFirst replacement backordered; correct part now in hand, kept for another day
C105Main2.2µF 400VYesPrimary side
C106Main47µF 25VYesUC2845 Vcc reservoir; replaced with Panasonic FR 47µF 50V, 0.27Ω ESR
C108, C109Main1200µF 200V 85°CYesThe bulk caps, designators hidden under the cans; the trace casualties happened here
C118-C121Main3900µF 10VNo5V output filters; verified healthy by ripple measurement
C127Main120µF 35VNo47µF 50V on JeffC’s board, so the cart part didn’t fit mine; correct part now in hand, kept for another day

The parts order

The complete bill of materials, for the next PEC4044B owner. It took two DigiKey orders. The first was JeffC’s cart. Note that the replacements carry higher voltage ratings than the originals where the modern part in the same footprint allows it (450V over 400V for C105, 35V over 25V for the small ones, 16V over 10V for the output filters).

Manufacturer partMakerValueQtyDestination
UVZ2W2R2MPDNichicon2.2µF 450V1C105
EEU-FR1H470BPanasonic (FR series)47µF 50V1C106
ESW476M035AE3AAKEMET47µF 35V2C236, plus a spare
ESC226M035AC3AAKEMET22µF 35V3C232, C253, C259
UHE1C472MHDNichicon4700µF 16V1C235
SLP122M200E4P3Cornell Dubilier1200µF 200V snap-in2C108, C109
UHE1H102MHD6Nichicon1000µF 50V1C233; backordered, never arrived
ELXZ160ELL392MK40SChemi-Con3900µF 16V4C118-C121; in the drawer, never installed

Subtotal: $35.97. That cart got me almost all the way there, but not quite: C127 on my board is 120µF 35V where JeffC’s is 47µF, and the 1000µF part for C233 was backordered. A second order finished the job.

Manufacturer partMakerValueQtyDestination
EEU-FC1V121BPanasonic (FC series)120µF 35V1C127
UPW1H102MHDNichicon1000µF 50V1C233

Subtotal: $2.96. Two orders, $38.93 all in. The supply is unobtainable. The complete bill of materials to save one costs less than forty dollars.

Where it stands

The recap stands deliberately unfinished: C118 through C121, C127, and C233 are still the originals, kept for a braver day. Everything else is done. The supply passed its bench validation, cold-start, warm-start, and hour-plus soaks, bodges and all, then went back into its case, and the case went back into the Indigo. The machine is fully validated and back in service. The photo at the top of this article is exactly what that looks like: the repaired supply running the machine into hour three.

Lessons

  • A supply that dies with rails in perfect regulation isn’t failing; it’s being told to stop. Scope the status/standby line, not just the outputs.
  • “Won’t restart until it sits unplugged a while” means a latch draining, not a part cooling.
  • Instruments lie in specific ways: a 100W dim-bulb starves a 300W supply, a once-a-second DMM invents phantom voltages, scan-mode scope displays alias, and ground-clip loops manufacture ripple. Every scary measurement deserves a control experiment.
  • Same model does not mean same board. Verify every cap value against your own silkscreen.
  • Old caps age individually. The 33-year-old that measured perfect and the three that (probably) killed the machine came off the same board, same brand, same date codes.
  • Control the blast radius. One board at a time, reassemble, retest, then move on. When something breaks, you know exactly where you last touched.
  • Meticulous diagrams still didn’t stop one backwards cap and one in the wrong hole. Inspection and in-circuit checks with the meter caught both before power did.
  • Knowing when to stop is a repair skill. Once the fault was cured, every further extraction was optional surgery on an irreplaceable board. The last caps can wait.

The cap map, parts order, and pin map are all above. The full day-by-day log is available if anyone’s fighting one of these. And for the next PEC4044B owner searching in vain: the bulk caps are C108 and C109. You’re welcome.