Forum Discussion
Floading Point divider, throughput/latency
Hi all,
I'm considering using the floading point blocks available from Altera in a design I have. Throughput is a top priority as I'm doing complex image transformation for high speed application. I was trying out the altfp_div megafunction and found out that it outputs only correct answers every other clock cycle, possibly worse as I only tried it with two numbers. Is the altfp_div not fully pipelined? I saw this also with latency set to 14. I have added figures with this where div6rest_result is the correct values and div6_result_real is the value form the altfp_div. I also did the same for the altfp_inv floating point inverter. It seems to output correct values when receiving data every cycle. I also added a zip file with this test made in Quartus 9.1 Cheers Stefan p.s. Second question I have is are those megafunctions free to use with the web edition?19 Replies
- Altera_Forum
Honored Contributor
--- Quote Start --- Hi nplttr, I'm fully aware of this. My experience is 3 years ASIC design (front-end and back-end design) and ~1 year FPGA design with Altera devices. In this example when I use the rising edge then the data has a full clock cycle to become stable before it is registered into the divider. How ever if I use falling edge the data only has half a cycle... I prefer to keep my design synced on rising edge unless I have a really good reason to switch over to falling edge. To place this buffer into the design will also cost my 64 registers + some additional logic. I wander if this is a bug in the behavioural code for the divider and if I have to do a gate level simulation for this... Cheers Stefan --- Quote End --- You don't have to change the edge of the clock that produces the data. It will work. If you want to have an RTL simulation that is consistent with the real device, let the input data change silightly before the edge of the clock that registers it. - Altera_Forum
Honored Contributor
As I previously checked, setting the data at the rising edge gives correct results in a timing simulation, e.g. using Quartus V9 simulator. That's what you also can expect in a real device, because FPGA registers have a zero hold time requirement.
The ModelSim results are with functional simulation however, which is apparently causing the problems. At first sight, I wanted to agree with Tricky: --- Quote Start --- All data in the testbench waits for the rising edge of the clock, so input should be safe. --- Quote End --- But it this actually true? The VHDL specifation guarantees, that a process is "executed" completely, before the signals are updated. The execution order of multiple processes is however undefined. Also we don't know, how altfp_div is organized internally. Possibly combinational logic is placed before the first register level. If the code is not well considered, simulation artefacts may occur, effectively creating pathes of different length in terms of simulation delta cycles up to the first register. This won't matter in synthesized logic, when actual LE delays apply. If it's so, the suggestion to set the data on falling edge, or generally a few simulation time steps away from active clock edge, will help. P.S.: --- Quote Start --- I wander if this is a bug in the behavioural code for the divider --- Quote End --- Yes, I suppose so. It should be avoidable by better considering delta cycle delays. - Altera_Forum
Honored Contributor
Hi nplttr,
I'm fully aware of this. My experience is 3 years ASIC design (front-end and back-end design) and ~1 year FPGA design with Altera devices. In this example when I use the rising edge then the data has a full clock cycle to become stable before it is registered into the divider. How ever if I use falling edge the data only has half a cycle... I prefer to keep my design synced on rising edge unless I have a really good reason to switch over to falling edge. To place this buffer into the design will also cost my 64 registers + some additional logic. I wander if this is a bug in the behavioural code for the divider and if I have to do a gate level simulation for this... Cheers Stefan - Altera_Forum
Honored Contributor
In a real system the input for a Flip Flop has to be stable for an interval of time before the active edge of the clock. This time is namd Setup Time (ts) and depends on the technology and the design of the Flip Flop.
In most simulations, it is safe to let the input data change on the falling edge of the clock, if the flip flop is triggered on the rising edge of the clock. In the real system, the input will be generated by another flip flop with combinational logic. The relation that has to be verified is the setup constraint that is: T > tq +Tc.max + ts Where: T = Clock period tq = flip flop delay ts = setup time Tc,max = the maximum combinational delay of the logic between the flip flops. - Altera_Forum
Honored Contributor
P.S. I forgot to say that I changed the testbench and now the data is changing on the falling edge of the clock.
- Altera_Forum
Honored Contributor
- Altera_Forum
Honored Contributor
There's something I don't understand.
I run a simulation of your circuit using the attached testbench. If I read weel you're doing the operation 10/512=0.19... and 512/10 =51.2. I used your function : div6_result_real <= float32ToReal(div6_result); to display the results and it looks correct. You can run the simulation and let me know. - Altera_Forum
Honored Contributor
All data in the testbench waits for the rising edge of the clock, so input should be safe.
- Altera_Forum
Honored Contributor
It looks like your input data is changing on the rising edge of the clock.
This is not safe. Can you try again ensuring that the input data is changing of the falling edge of the clock?