Forum Discussion
Aria V Display Port core sometimes disconnected
Hi All
We have a PCB board we designed with AMD 8860 GPU that sends DP video to Aria V Display Port core (Qsys). In some of our boards we sometimes have DP connection issues.
The video signal goes to altera_xcvr_native_av and when RX is out of lock we see multiple retries on the DP AUX channel.
We came across this bug:
in our case the enable GPU is not checked.
Can anyone confirm this issue is also applicable to Aria V Display Port?
Thanks
Ariel
50 Replies
- relsaar
New Contributor
Hi dlim
I'd like to test our findings by comparing the display port BER between 2 boards. a board that tends to fail with a board that is usually stable.
1. Is there a signal that I can add to the SignalTap to give numerical value for BER or shall I accumulate disp_err or code_err toggles per time unite?
2. What is the threshold for a reasonable BER?- relsaar
New Contributor
Hi,
I checked disp_err and code_err signals on a system that never had msa_lock toggles and saw the same toggles on those signals (similar to the photo I attached before)
also we only use one lane of display port but in the signalTap all 4 of them are noisy. it looks like we can't trust this BER detection/debug signal.
I tried to connect disp_err or code_err as enable of a counter to quantify the BER and check if the the first board has more toggles on that indicates higher BER but unfortunately those signals comes from encrypted module and can't be used.
thanks
Ariel Saar
- Deshi_Intel
Regular Contributor
MSA log dump example
- Deshi_Intel
Regular Contributor
Hi,
Since you are using altera_dp IP which means this is Altera DisplayPort IP, not Bitec DisplayPort IP.
- You should be able to upgrade to latest Quartus version after you sort out the licensing issue with your local dFAE support. Pls try out Quartus Standard v20.1 if possible
- The reason you are seeing a lot of Bitec design internally is due to Altera previously purchased Bitec DisplayPort IP, modified and repackage into Altera DisplayPort IP
Now while you are reviewing your board design, let's sort out the debug step to capture the right debug status signal first in order to determine the right debug direction.
- DP Rx data flow from GPU -> FPGA (FPGA Transceiver Rx channel -> DP Rx Sink IP)
- That's why I mentioned earlier if Transceiver Rx channel can't capture/sample the Rx data properly then the data will break/corrupt in transceiver channel before reaching DP Rx IP. No point to debug DP RX IP anymore.
- We can monitor "Rx_is_lockedtodata" signal to check whether Rx channel CDR is in locked mode or not. Ignore "Rx_is_lockedtoref" as you are not using it. Both are status signal of CDR but it's meant for different CDR operation mode. You should be referring to "Rx_is_lockedtodata"
- CDR lock indicated the data rate speed is correct and Rx channel can sample data at the correct data rate speed. (but we won't know the correct data content is receive or not)
- So, can you tell me whether "Rx_is_lockedtodata" stay high or de-asserted in your video test failure scenario ? Pls share with me the signal_tap file result
- Then next is to check whether correct video data content is sample correctly at DP Rx IP or not. You can monitor via MSA lock signal (which you already did)
- You mentioned MSA loose lock. One of the possibility is due to data corruption due to high "bit error rate (BER)", bad signal integrity on the board - either on the refclk source or the DP Rx channel itself.
- We can monitor DP soft 8b/10B decoder IP status signal like *code_err" and "disp_err" to verify whether it's related to high BER or not. Can you try to find these status signal in your signal_tap and capture them in failure condition ?
- Another possibility that I can think of is there is limitation of DP bandwidth transfer.
- From your DP Rx sink IP setting, you configure it to 1.62G with 1 lane count only
- May I know what exact video resolution and bit per colour that you set in your GPU ?
- Attached is the DP BW calculation guideline to help you verify you video data transfer doesn't exceed DP link BW
- Another thing I noticed is your "pixel output mode" = single.
- This may put some pressure on your design timing closure as you need higher DP operating frequency
- May I know is your design timing closed at Quartus Timequest ?
- Also, maybe you can consider to change the setting to "quad" to relax the design timing requirement
Thanks.
Regards,
dlim
- relsaar
New Contributor
Hi dlim
I was able to recreate the msa_lock fail and during this event Rx_is_lockedtodata looks stable (attached SignalTap image). I guess the problem is not the ref clock.
As you suggested I monitored DP soft 8b/10B decoder IP status signal like “code_err" and "disp_err" and those signals look very noisy even during normal operation
(attached SignalTap image). During failure those error signal looks the same.
Can we assume the problem is due to display port poor signal integrity?
Regarding your inquire for video resolution and bit per color that was set in our GPU, please see calculation below:
720 X 576 X 25Hz X 16 bps = 165,888,000 bps
I assume this is a very low channel capacity since we have 1.62 GHz bit per second hence the "pixel output mode" = single is not an issue. I think we tried using "quad" and it complicated the clocks in the design. Let me know if you disagree my assumptions.
We don’t have any special constrains for those IPs, only defaults that were applied after we added the IP. However after I ran Report Top Failing paths I got:
qsta_utility::generate_top_failures_per_clock "Top Failing Paths" 200
No failing paths found
Thanks
Best Regards,
Ariel Saar
FPGA Design Engineer
- Deshi_Intel
Regular Contributor
Hi,
I think the first thing we need to clarify is which DisplayPort IP that you are using here ? Are you using Bitec DisplayPort IP or Altera DisplayPort IP ?
- If you are using Bitec DisplayPort IP
- Then my suggestion to you will be to engage back Bitec for better issue support
- I also share same concern with you on IP upgrade capability. Since this is not Altera IP, Intel Engineering won't support IP migration to latest Quartus version. Bitec should has their own IP migration strategy. Pls consult Bitec accordingly
- Else if you are using Altera DisplayPort IP
- Then it only make sense for you to perform IP migration to latest Quartus version.
- Byright evaluation license should still able to generate sof file but I am not familiar with licensing issue. Perhaps you can contact your local FAE for licensing support ?
I presume you are referring to Rx channel CDR lock monitor signal (rx_is_lockedtodata)
- Attached is the CDR loose lock debug checklist for your reference. It's targeted for C10 GX FPGA but the debug methodology is more or less the same
- If you confirmed CDR is indeed loose lock on your board then there is no point to dump MSA log to debug DisplayPort anymore because your Rx transceiver channel already failed to sample/receive data in the first place before the data is passed to DisplayPort RX IP
- When PPM threshold setting = 1000, it means it allow max PPM drift until 1000ppm while your oscillator ppm is only +-20ppm which is good. So, maybe ppm is not the issue here
- Yup, your 135MHz refclk signal probe result looks bad but I am not sure is it due to your board design issue, or your probing contact issue or oscilloscope bandwidth issue ?
- Did you manually hold your diff probe or use some tool to hold it or solder the probe tip on board ?
- Also is your oscilloscope bandwidth at least 3x or 5x of 135MHz ?
- Lastly, yes. Enable 100ohm OCT insides FPGA should help to improve the LVDS signal termination. Pls try it out.
- You also mentioned your 1.67G Rx data transfer signal quality should be good
- May I know how do you check it ? Did you run board simulation or use transceiver toolkit to to perform loopback testing on board ?
Thanks.
Regards,
dlim
- relsaar
New Contributor
Hi dlim
We are using the IP catalog core from Quartos 15.0, I guess it is Intel Altera core but many of the generated files contain the word "bitech" on it. (see attachment). We intend to meet the local FAE tomorrow and check this issue.
Regarding the lock signal I mentioned earlier - I'm sorry for the confusion - I was referring the MSA - lock that can be generated on the QSYS GUI as a trigger(see attached image)
My signalTap also contain the below Transceiver Native PHY signals:
- Rx_is_lockedtoref
- Rx_is_lockedtordata
I assume the first is indicating on ref clock signal integrity and the second on display port data signal integrity?
Regarding the signal integrity issue,
- Scope and probe BW are 4GHz
- Probe was soldered to tips on board and was clumped to the PCB.
- 100ohm OCT was not helpful
Regarding your question on how we verify the video – we output the digital video signal from the Aria V in to external video encoder which convert it to analog screen. I was looking for the transceiver toolkit feature you mentioned but didn’t find such option. Can you please send me more info on how to use this feature?
However after I reviewed the brd file I found out the LVDS clock pins, though short, have no ground plan above or below the clock signals. I suspect it is the root cause for the noise. In order to lower the clock noise I replaced the 162MHz clk with 50MHz clk and use Alt PLL to convert the clock back to 135MHz. It looks like the signal is less noisy now (see attached image).
I’m running some tests to see if the failure is resolved or can be recreated again. Will update as soon as I have some answers.
- Deshi_Intel
Regular Contributor
- If you are using Bitec DisplayPort IP
- Deshi_Intel
Regular Contributor
Hi Ariel,
I looked at the KDB link. You can see at the top, it mentioned the affected FPGA family is just Arria 10 and Stratix 10 which means Arria V is not affected by this issue.
- Anyway, I always recommend user to upgrade to latest Quartus Standard edition version if possible to mitigate the risk of known issue from older Quartus version.
Btw, you mentioned Rx loose lock. Are you referring to Rx channel CDR locked to data signal or some other status signal ?
- Common factor that will caused CDR to loose lock are like below
- high PPM different on CDR refclk - checked your on board crystal or clock generator quality.
- Bad signal integrity on CDR refclk - same debug approach as above
- Bad signal integrity on data transfer on Rx channel - reduce your video resolution to reduce the DP bandwidth from 5,4G to 2.7G or even 1.67G to help isolate signal integrity issue
- You can also use DP example design to dump the MSA log to check if you observe high bit error rate (BER) on the 4 DP channel or 2 DP channel
Thanks.
Regards,
dlim
- relsaar
New Contributor
Hi dlim
Thank you for your fast and detailed response.
I understand the Arria V and 10 Display Port IP core is different and the bug report I mentioned earlier is not relevant for our problem. (please confirm)
Regarding your suggestion to migrate to a new version of Quartos - our Display port core was bought from Bitech (which was later bought by Altera) and is not supported above Quartos 16.0.
We tried to compile on Quartos 20.1 the reference design of current display port core but could not get an evaluation *.sof file. If migrating will solve the issue we will be happy to purchase a license or maintenance.
the lock signal I mentioned earlier is the output of the Arria V Transceiver Native PHY. The setting of this core has a PPM detector threshold that is set to 1000 PPM (similar to reference design) while our reference clock Oscillator is +/- 20PPM (see attached spec). I figure high value on the PPM detector makes the design more robust?
Our video resolution BW is already 1.62G and the signal looks good on debug monitor while it is locked, however the reference clock seems a bit noisy when we probe it with a deferential probe (see attached image). do you recommend to add "Input Termination" at "OCT 100 Ohms" in the assignment editor? Our differential ref clock pin is currently set to I/O Standard LVDS. Please see clk ref schematics image attached below, the right side nets "DP REF CLK+" and "DP REF CLK-" are connected directly to the Arria V FPGA.
We tried to check the reference design for dumping the MSA log but it is not clear how to apply it (Please send link or more info on this issue)
Thanks
Ariel
- Deshi_Intel
Regular Contributor
- Alberto_Sykes
Occasional Contributor
relsaar, Thank you for posting in the Intel® Communities Support.
In reference to your inquiry, in order for us to be able to provide the most accurate assistance for this scenario, I will transfer your thread to the proper department, they will further assist you with this matter.
Regards,
Albert R.
Intel Customer Support Technician
A Contingent Worker at Intel