Forum Discussion
Hardware VS Software
I am doing some performance analysis for hardware and software. The program will generate a set of random 32 bits number. these numbers will be send to a user peripheral (hardware) for addition. then, these numbers will be used again for addition using ordinary C code (software). this is the result: --Performance Counter Report-- Total Time: 0.0771603 seconds (3858017 clock-cycles) +---------------+-----+-----------+---------------+-----------+ | Section | % | Time (sec)| Time (clocks)|Occurrences| +---------------+-----+-----------+---------------+-----------+ |Hardware | 50.7| 0.03911| 1955390| 1| +---------------+-----+-----------+---------------+-----------+ |Software | 49.3| 0.03805| 1902555| 1| +---------------+-----+-----------+---------------+-----------+ suppose that I guess hardware should be faster but it turned out that software to be faster. may i know why?
20 Replies
- Altera_Forum
Honored Contributor
ok thanks... because last time when i was doing my lab, multiplication doesnt work, so i change to addition... so in my case, random 32-bit number should work fine for multiplication, right?
- Altera_Forum
Honored Contributor
integers can be multiplied without a library, they are part of the standard types.
unsigned and signed are in the ieee.numeric_std package. - Altera_Forum
Honored Contributor
how about the library?
- Altera_Forum
Honored Contributor
yes it can.
c <= a * b; is perfectly acceptable. As long as you do it with a type that has a multiply function (eg. integer, unsigned or signed) - Altera_Forum
Honored Contributor
how to do multiplication in VHDL? seems like the symbol "*" cannot be used...
- Altera_Forum
Honored Contributor
Adding is hardly an issue in software when you're within 32 bits. But when you look at larger bit type (like > 64) or fixed point you would start to see a difference. Then multiply that up by large data sets (eg. video or mega pixel images) you would really notice the difference in throughput.
- Altera_Forum
Honored Contributor
--- Quote Start --- I would really like to see the C code that adds the numbers so I can compare the software cycle count that you have with the cycle count that I get running the same code on my processor that directly executes the C code without compiling to a native instruction set. Will you please attach the code to this thread? The whole object of the design is to minimize the number of cycles so it fits right in with what you are doing. --- Quote End --- the code is like you generate 2 sets of random number, then you just add it up. for the "sub processor" that do the adding, you need to have 3 submodules: adding module, interface and the top level system. adding module is where you add up the generated number, interface is controlling the input and output and top level system is like whole system which includes interface and adding module. - Altera_Forum
Honored Contributor
--- Quote Start --- If you are shipping really short amounts of work to the accelerator then yes software will be faster due to the communication overhead. The only way this could be efficient is if you perform the same operation across a large block of data in memory and you use DMAs to stuff the data into the accelerator. --- Quote End --- Yes, my project is to test the performance of the system with and without using DMA. But for the basic one, I try to test the performance between hardware and software first without involving dma. next step, i will try to include dma. by the way, is there any tutorial regarding how to transmit and receive data for dma in c language? - Altera_Forum
Honored Contributor
I would really like to see the C code that adds the numbers so I can compare the software cycle count that you have with the cycle count that I get running the same code on my processor that directly executes the C code without compiling to a native instruction set. Will you please attach the code to this thread? The whole object of the design is to minimize the number of cycles so it fits right in with what you are doing.
- Altera_Forum
Honored Contributor
If you are shipping really short amounts of work to the accelerator then yes software will be faster due to the communication overhead. The only way this could be efficient is if you perform the same operation across a large block of data in memory and you use DMAs to stuff the data into the accelerator.