Forum Discussion

Altera_Forum's avatar
Altera_Forum
Icon for Honored Contributor rankHonored Contributor
15 years ago

Floading Point divider, throughput/latency

Hi all,

I'm considering using the floading point blocks available from Altera in a design I have. Throughput is a top priority as I'm doing complex image transformation for high speed application.

I was trying out the altfp_div megafunction and found out that it outputs only correct answers every other clock cycle, possibly worse as I only tried it with two numbers. Is the altfp_div not fully pipelined? I saw this also with latency set to 14.

I have added figures with this where div6rest_result is the correct values and div6_result_real is the value form the altfp_div.

I also did the same for the altfp_inv floating point inverter. It seems to output correct values when receiving data every cycle.

I also added a zip file with this test made in Quartus 9.1

Cheers

Stefan

p.s. Second question I have is are those megafunctions free to use with the web edition?

19 Replies

  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    --- Quote Start ---

    Hi nplttr,

    I'm fully aware of this. My experience is 3 years ASIC design (front-end and back-end design) and ~1 year FPGA design with Altera devices.

    In this example when I use the rising edge then the data has a full clock cycle to become stable before it is registered into the divider.

    How ever if I use falling edge the data only has half a cycle...

    I prefer to keep my design synced on rising edge unless I have a really good reason to switch over to falling edge. To place this buffer into the design will also cost my 64 registers + some additional logic.

    I wander if this is a bug in the behavioural code for the divider and if I have to do a gate level simulation for this...

    Cheers

    Stefan

    --- Quote End ---

    You don't have to change the edge of the clock that produces the data.

    It will work.

    If you want to have an RTL simulation that is consistent with the real device, let the input data change silightly before the edge of the clock that registers it.
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    As I previously checked, setting the data at the rising edge gives correct results in a timing simulation, e.g. using Quartus V9 simulator. That's what you also can expect in a real device, because FPGA registers have a zero hold time requirement.

    The ModelSim results are with functional simulation however, which is apparently causing the problems. At first sight, I wanted to agree with Tricky:

    --- Quote Start ---

    All data in the testbench waits for the rising edge of the clock, so input should be safe.

    --- Quote End ---

    But it this actually true? The VHDL specifation guarantees, that a process is "executed" completely, before the signals are updated. The execution order of multiple processes is however undefined. Also we don't know, how altfp_div is organized internally. Possibly combinational logic is placed before the first register level. If the code is not well considered, simulation artefacts may occur, effectively creating pathes of different length in terms of simulation delta cycles up to the first register. This won't matter in synthesized logic, when actual LE delays apply.

    If it's so, the suggestion to set the data on falling edge, or generally a few simulation time steps away from active clock edge, will help.

    P.S.:

    --- Quote Start ---

    I wander if this is a bug in the behavioural code for the divider

    --- Quote End ---

    Yes, I suppose so. It should be avoidable by better considering delta cycle delays.
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    Hi nplttr,

    I'm fully aware of this. My experience is 3 years ASIC design (front-end and back-end design) and ~1 year FPGA design with Altera devices.

    In this example when I use the rising edge then the data has a full clock cycle to become stable before it is registered into the divider.

    How ever if I use falling edge the data only has half a cycle...

    I prefer to keep my design synced on rising edge unless I have a really good reason to switch over to falling edge. To place this buffer into the design will also cost my 64 registers + some additional logic.

    I wander if this is a bug in the behavioural code for the divider and if I have to do a gate level simulation for this...

    Cheers

    Stefan
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    In a real system the input for a Flip Flop has to be stable for an interval of time before the active edge of the clock. This time is namd Setup Time (ts) and depends on the technology and the design of the Flip Flop.

    In most simulations, it is safe to let the input data change on the falling edge of the clock, if the flip flop is triggered on the rising edge of the clock.

    In the real system, the input will be generated by another flip flop with combinational logic.

    The relation that has to be verified is the setup constraint that is:

    T > tq +Tc.max + ts

    Where:

    T = Clock period

    tq = flip flop delay

    ts = setup time

    Tc,max = the maximum combinational delay of the logic between the flip flops.
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    P.S. I forgot to say that I changed the testbench and now the data is changing on the falling edge of the clock.

  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    I also attach a print of the screen.

    The latency is 75.76ns.
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    There's something I don't understand.

    I run a simulation of your circuit using the attached testbench.

    If I read weel you're doing the operation 10/512=0.19... and 512/10 =51.2.

    I used your function :

    div6_result_real <= float32ToReal(div6_result);

    to display the results and it looks correct.

    You can run the simulation and let me know.
  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    All data in the testbench waits for the rising edge of the clock, so input should be safe.

  • Altera_Forum's avatar
    Altera_Forum
    Icon for Honored Contributor rankHonored Contributor

    It looks like your input data is changing on the rising edge of the clock.

    This is not safe.

    Can you try again ensuring that the input data is changing of the falling edge of the clock?