Forum Discussion
NiosII is slow? Or something I am doing wrong.
Hi all,
I am playing with Cyclone II chip. It is custom board and there is ATMEGA16 running at 12MHz and FPGA with master clock 50MHz. FPGA is translating data from internal dualport memory via LVDS to other board with lots of RGB LEDs. Just for fun and for learning purposes I wrote small program for Mandelbrot fractal calculation in C. First of all, I compiled it with GNU C for ATMEGA. When I copy-pasted same code to Nios II. The difference between tests: a) when runing from ATMEGA, CPU clock is 12MHz, data passed to FPGA via 4 wires + control line. Both nibles are joined in FPGA and RAM adr is calculated in hardware. b) when runing from NIOS II, CPU clock is 50MHz, data passed to RAM in 8 wires, RAM adr is passed via 12 wires directly. No need for calculation. Nios II is configured with hardware float multiplication. All RAM is on-chip. Guess who is faster? I didn't made exact measurement. But BOTH system looks like runing on same speed! :eek: What is wrong? Testing software: #include "main.h" static float zeta;   void pushbyte(char x, int adr) { IOWR_ALTERA_AVALON_PIO_DATA(RAMADR_BASE,adr); IOWR_ALTERA_AVALON_PIO_DATA(PORTB_BASE, x); IOWR_ALTERA_AVALON_PIO_DATA(PORTA_BASE, 1); IOWR_ALTERA_AVALON_PIO_DATA(PORTA_BASE, 0); }     void pumprgb(void) { int adr; float cx, cy, scale, a1, b1, a2, b2, ax, ay; signedint x,y; long color; float a12, b12; int limit; int lp; cx=-1.52; cy=0; scale=0.05001-(zeta*0.002); limit=4; if (zeta>128) {zeta=0;} //sitas netelpa! 100 baitu reikia.   y=-12; while (y<12) { x=-40; ay=cy+y*scale; while(x<40) { ax=cx+x*scale;   b1=ay; a1=ax; a12=a1*a1; b12=b1*b1; lp=0; while ((lp<255) && ((a12+b12)<limit)) { lp++; a12=a1*a1; b12=b1*b1;   a2=a12-b12+ax; b2=2*a1*b1+ay; a1=a2; b1=b2; } color=lp; color=color*45536; adr=((int) ((x+40)*3+((y+12)*240))) & 0x1FFF; //adr=adr & 0x1FFF; pushbyte((color>>16),adr); // red pushbyte((color>>8),adr+1); // green pushbyte(color,adr+2); //blue x++; } y++; } zeta++; }   int main(void) { zeta=0; while(1) { pumpRGB(); } return(0); }14 Replies
- Altera_Forum
Honored Contributor
Here is video of this test: http://www.youtube.com/watch?v=7yscdt3lo10
There is small bug in ATMEGA version- Cyclone configuration is still not loaded but Atmega is already running. Next frame is completely synchronized. So image jumps. But lets concentrate at the speed. - Altera_Forum
Honored Contributor
Are you compiling everything with -O2 or -O3 ?
The default for the IDE is -O0, and that will generate atrocious code. - Altera_Forum
Honored Contributor
I think the II/f looks significantly faster. The II/e core is _not_ designed for any sort of performance application. It's silly to even attempt it if you know what's happening...one instruction every 6 clock cycles.
Besides, you don't use a softcore for performance reasons. You use it "because it (or the FPGA) is there" or because you want to take advantage of FPGA hardware (RTL) acceleration for your algorithms/protocols/whatever. Cheers, --slacker - Altera_Forum
Honored Contributor
--- Quote Start --- Are you compiling everything with -O2 or -O3 ? The default for the IDE is -O0, and that will generate atrocious code. --- Quote End --- I am using default configuration. I'll try to change these options. - Altera_Forum
Honored Contributor
--- Quote Start --- Besides, you don't use a softcore for performance reasons. You use it "because it (or the FPGA) is there" ... --- Quote End --- There is full emulation of A500 computer on DE1 bord and softcore 68K is quite fast. But first I'll try optimisation options as in other reply. - Altera_Forum
Honored Contributor
I would expect this algorithm to be slow. This algorithm is performing many floating point calculations and without the floating point custom instruction all of that would end up being emulated in software using Nios II. The algorithm is also inherently slow as are most convergence algorithms and as a result you can speed it up by breaking down the algorithm and handling it in parallel (multi-core or hardware accelerators).
Why the performance is different between the two cores I'm not sure but make sure you are using the same compiler optimizations so that you are comparing apples to apples. - Altera_Forum
Honored Contributor
--- Quote Start --- I would expect this algorithm to be slow. This algorithm is performing many floating point calculations and without the floating point custom instruction all of that would end up being emulated in software using Nios II. The algorithm is also inherently slow as are most convergence algorithms and as a result you can speed it up by breaking down the algorithm and handling it in parallel (multi-core or hardware accelerators). Why the performance is different between the two cores I'm not sure but make sure you are using the same compiler optimizations so that you are comparing apples to apples. --- Quote End --- Maybe you didn't understand main post. English is not native language for me. I want to emphasize, that ATMEGA16 do not have any FP instruction at all. NiosII in this experiment has hardware multiplication, addition and substraction (software emultation didn't fit to this fpga). Nios is 32 bits wide (how many bits floats have?) and Mega is only 8 bit. And clock speed of Nios is 4 times! faster. I do not modify algorithms as it is fair play- same code is copy-pasted from one platform to other. And the question "Why the performance is different between the two cores" is answered in very simple way. Who will pay money for "f" version if "e" version is same. :) - Altera_Forum
Honored Contributor
Watching the video you posted, it looks to me that the Nios versions are faster than the ATmega version. By the time the ATmega has finished one frame, the Nios II/e has finished almost 3 and the Nios II/e has finished almost 5 frames.
Given the 32 bit vs 8 bit and floating point vs non floating point differences, I too would have guessed there would be a bigger difference. You would have to analyze the compiled code of each in detail to see why it performs the way it does. You still have not confirmed you actually enabled optimizations in the Nios II version. As others have pointed out -O0 is the default optimization setting which will produce very inefficient code. Another though I had is that you may want to verify your Nios code is actually using the FP hardware. I don't have any experience with this, but I have seen quite a few other posts on these forums about people having trouble with the compiler not actually using the FP hardware but using software emulation instead. - Altera_Forum
Honored Contributor
--- Quote Start --- Another though I had is that you may want to verify your Nios code is actually using the FP hardware. I don't have any experience with this, but I have seen quite a few other posts on these forums about people having trouble with the compiler not actually using the FP hardware but using software emulation instead. --- Quote End --- It is very simple: software emulation of floating point will not fit to this Cyclone device's internal RAM. It is very small device and I must leave some dualport "video" RAM for LVDS data transfer. And about optimization. I checked config, and it is -O2. Also, small C, no software floating point, no clean exit and etc. I need to strugle for each byte of RAM. - Altera_Forum
Honored Contributor
I got curious so I compiled the code you posted for both Nios and for AVR and compared the compiled code.
One thing I noticed that may account for the lack of expected performance is that Nios does not use any hardware acceleration for FP comparisons. For example, your inner loop does a comparison with the limit variable. On Nios this calls a subroutines totaling about 183 opcodes. (Less than this will actually execute due to the different possible branches taken, but this is a quick number for comparison purposes.) I also looked at the AVR code for the same comparison. If you link in the hand optimized assembler libm.a library, there are a total of about 40 opcodes. This is a lot more efficient than the Nios case. If you don't use libm.a but instead let GCC use it's own library (as is done with Nios), then the total opcodes is a little over 400. Are you using libm.a with your AVR tests? If you are just trying to compare performance of the architectures, a more fair test would be to not use libm.a with AVR. I have a feeling the lack of performance is mostly due to the poorly optimized FP operations for which Nios does not have hardware acceleration. I wonder if there is an easy way to accelerate the FP comparisons with Nios? It would probably make a big difference if you did this. Another similar inefficiency I noticed is the conversion of x, y and limit to floating point values. I suspect even changing the compare to limit to a compare to a constant value of 4 will make a small noticeable difference. The compiler is not optimizing the fact that limit is actually a constant.