Forum Discussion
Cyclone IV GX Oscillator failures
- 2 years ago
1) 50 MHz oscillator with 50 MHz design in the FPGA was tested w/ no failures. 10 of these units are being used in the field.
2) During qualification testing, it's discovered that an 8 MHz output is actually 8.006 MHz. The internal divider is set to 32/25 and the output should be exact. I couldn't fix this in the FPGA.
3) I change the 50 MHz oscillator to 64 MHz and the PLL to 1/1 and now my 8 MHz is really 8 +/- 50 ppm.
4) All of the boards that I have are re-tested and a few are sent to be used.
5) One of the boards that I have and one in the field show no activity on the PCIe. Then 2 more.
6) The PLL that drives the PCIe still has an INCLK value of 50 MHz. The downstream values of the PLL should be 50, 125, 312.50, and 1250 but are actually running at 64, 160, 400, and 1600. I fixed the INCLK value of that PLL, but the boards do not recover.
7) The failed boards have little or no amplitude on the 64 MHz, so we replace those oscillators. The oscillators all fail within 24 hours.
We bought several different kinds of oscillators -- same results. -- The oscillators all fail within 24 hours. 9) I replaced the 22 ohm series termination with 220 -- same results. -- The oscillators all fail within 24 hours.
10) I contacted EPSON and Abracon engineers who told me that the oscillators will tolerate a short circuit indefinitely and recover when the short is removed.
11) I replaced 2 FPGAs so far and they have been in continuous operation for 3 days now. I'm sending 2 more out to be replaced.
12) I have 11 boards that have never failed in a burn-in environment and have been running for 3 days.
So, YES the FPGAs were damaged and somehow killed bulletproof oscillators along the way.
In the timing analyzer, the higher clock values are shown to be derived. Other than that, there was no way to tell that my chips were time bombs in the field.
what is "%28 change in INCLK value" in commonly understood technical terms? You are talking about frequency, voltage?
a 50 mhz oscillator was replaced with a 64 mhz before re-programming the pll and the FPGA was permanently damaged.
- _AK6DN_2 years ago
Frequent Contributor
"a 50 mhz oscillator was replaced with a 64 mhz before re-programming the pll and the FPGA was permanently damaged".
What?
In your first post you said:
"Replacing the oscillator fixes the board temporarily. The oscillators fail again several hours to a day later. We have tried several different types and manufacturers of the oscillators. "
So which is it? Bad FPGA or bad oscillators? Or something else?
50MHz and 64MHz are both well within the input spec of the FPGA clock input.
I still don't think you have root caused the failure.
- bob_bitchen2 years ago
Occasional Contributor
1) 50 MHz oscillator with 50 MHz design in the FPGA was tested w/ no failures. 10 of these units are being used in the field.
2) During qualification testing, it's discovered that an 8 MHz output is actually 8.006 MHz. The internal divider is set to 32/25 and the output should be exact. I couldn't fix this in the FPGA.
3) I change the 50 MHz oscillator to 64 MHz and the PLL to 1/1 and now my 8 MHz is really 8 +/- 50 ppm.
4) All of the boards that I have are re-tested and a few are sent to be used.
5) One of the boards that I have and one in the field show no activity on the PCIe. Then 2 more.
6) The PLL that drives the PCIe still has an INCLK value of 50 MHz. The downstream values of the PLL should be 50, 125, 312.50, and 1250 but are actually running at 64, 160, 400, and 1600. I fixed the INCLK value of that PLL, but the boards do not recover.
7) The failed boards have little or no amplitude on the 64 MHz, so we replace those oscillators. The oscillators all fail within 24 hours.
We bought several different kinds of oscillators -- same results. -- The oscillators all fail within 24 hours. 9) I replaced the 22 ohm series termination with 220 -- same results. -- The oscillators all fail within 24 hours.
10) I contacted EPSON and Abracon engineers who told me that the oscillators will tolerate a short circuit indefinitely and recover when the short is removed.
11) I replaced 2 FPGAs so far and they have been in continuous operation for 3 days now. I'm sending 2 more out to be replaced.
12) I have 11 boards that have never failed in a burn-in environment and have been running for 3 days.
So, YES the FPGAs were damaged and somehow killed bulletproof oscillators along the way.
In the timing analyzer, the higher clock values are shown to be derived. Other than that, there was no way to tell that my chips were time bombs in the field.
- bob_bitchen2 years ago
Occasional Contributor
It's bad FPGAs.
The FPGAs failed in a way that was surprising.
In the original design, the 50 MHz clock drove a qsys pll that used 32/25 as a divider to produce an 8 MHz output. This divider said that it was exact in the GUI, however it was off by about 1/1000.
In the second iteration of the design the 64 MHz clock used a divider of 1 and the 8 MHz output was correct to 50 ppm. The PCIe pll was supposed to produce 50, 125, 312.50 and 1250 outputs. They were actually 64, 160, 400, and 1600, and the PCIe was functioning.
There were no indications of a problem unless you looked at the derived clocks from the timing analyzer and noticed the values to be wrong.
In the third design, the INCLK of the PCIe pll was fixed to be 64.
I communicated with engineers from EPSON and Abricon and both told me that the devices are capable of driving a short circuit indefinitely and then fully recovering when the short is removed.
We have replaced 2 FPGAs and they have been operating for 3 days each. I sent 2 more out to be replaced.
I have also put a handful of boards that have never failed into a burn-in environment, and they have been in operation for the same amount of time.
Somehow, the FPGAs have managed to damage bulletproof oscillators while failing in a very obscure way.
- _AK6DN_2 years ago
Frequent Contributor
My reading of the timeline is that you compromised or damaged either the FPGA or OSCILLATOR or both during the first rework process to replace the 50MHz parts with 64MHz parts. The initial boards (built by a PCB assembly vendor I expect) worked 100%.
Only after rework of the 50MHz to 64MHz parts did failures occur.
My conclusion would be damage during that rework, could be an ESD event, high temperature, or something related. I am guessing the rework was done by hand by a local tech.
Subsequent rework replacing oscillators and FPGAs was OK. I am guessing that rework was done by a specialty rework tech, as BGA rework requires special equipment and skills.