Forum Discussion
Code in tightly coupled memory
Hi all,
I placed a function in a section different from the main code. More exactly, my code is mapped to sdram and I linked a function testcall() into a tightly coupled memory section which I previoulsy defined in sopc builder. So I declared: int testcall(int n) __attribute__ ((section (".tc_code"))); Then, following the suggestion I found in another thread I called ALT_LOAD_SECTION_BY_NAME(tc_code) before calling testcall(). My questions are: - is the ALT_LOAD_SECTION_BY_NAME(tc_code) mandatory? The code seems to work even if I don't use it. - if tc_code section is located in internal fpga memory (cyclone III M9k blocks), how much is the code speed improvement I could roughly obtain with respect to sdram (assume I'm using 2 tse with sgdma). Thank you Cris37 Replies
- Altera_Forum
Honored Contributor
I have several functions and many variables in onchip RAM and have no problems with any code or data I moved into it. The speed increase over SDRAM is very substantial. In fact, it can be too fast - I need interpacket delays sending a UDP data stream to the PC or I overrun my 3+GHz quad-core PC.
I have 10k of onchip RAM for code and data so a not insignificant portion of my program is stored there. I did not use ALT_LOAD_SECTION_BY_NAME. Bill - Altera_Forum
Honored Contributor
Cris, I tested your function in an example design on my Altera StratixII Devkit.
(C:\altera\91\nios2eds\examples\vhdl\niosII_stratixII_2s60_RoHS\TSE_SGDMA with the hello_ucosii software example) I have done some minor changes in your code and putted them into the example. There are two tasks which are called periodically. Before they initiated by the OS the function testcall() is called. Then every time when task2 is executed the function pkt_send() is called. Both works fine. Please check the attached projects. Is this that what you wanted to do? Jens - Altera_Forum
Honored Contributor
@BillA
I'm also sending (hardware generated) UDP packets to a PC. The best results are achieved if I have a point to point connection from hardware to PC with a separate network adapter just for the hardware. For some Intel NICs (Pro 1000 series) you can increase the number of receive descriptors (see attached jpg). This reduces in my case significant loss of packets. Jens - Altera_Forum
Honored Contributor
Thank you Jens,
your sample is basically what I do. The only difference being that I have all the other application stuff. Nevertheless I can't make mine work! Latest tests I performed: - same behaviour if these functions are mapped to tcm, sram or any other section different from sdram where main code is stored - whenever I compile in debug mode, none of the function works; in release mode usually testcall is working. - I also have a tc_data section, similar tightly coupled memory as tc_code, but for data. I used it both for system stack and some app data and I had no problem. - Altera_Forum
Honored Contributor
Hmm, that's tricky. Did you move the __attribute__ stuff into the header? Next you can try to unload the cache before (or after?) you call the function. Could you test execution of code from the other memories by changing the .text (heap, stack) section mapping in system libraries properties? There can be many other reasons ... task stack size to small?
I think I would start with reduced functionality in hardware and software. (like in the example). Then try to extend it step by step. Jens - Altera_Forum
Honored Contributor
I discovered that code is NOT actually loaded into tc_code section (all reads 0xff)
The small testcall works in release mode because the optimizer inlines it in the caller. The bigger pkt_send function, on the other hand, is really in tc_code. In debug mode there is no call optimization, so both function don't work. The behavior is actually very strange: the linker map file tells me that testcall has been placed in tc_code but if I step the assembly code with the debugger I see it executes as if it is inlined into the caller. !?!??!?? So, the ultimate question is: how can I force to load the tc_code (or sram, or whatever) section? ALT_LOAD_SECTION_BY_NAME(tc_code) is useless. Probably tc_code is ignored because in syslib properties .text is mapped to sdram and then only this memory is loaded? - Altera_Forum
Honored Contributor
I have attached the objdump file from the example. There you can see that the functions are mapped into tc_code (debug mode).
You can enable generating objdump files in the Nios IDE (Window->Preferences->NiosII->create objdump file) The syslib properties does not have effect to the __attribute__ directive. Could you try to execute a small programm from any other RAM than SDRAM to check if tc_code onchip RAM (SSRAM, ...) is working correct? Then you must change the system lib properties. Jens - Altera_Forum
Honored Contributor
Try adding __attribute__((noinline)) to the function prototype. As in:
That will force a function call. You might have the code body existing for any external calls - with the local call being inlined.static void foo(void) __attribute__((noinline)); static void foo(void) { .... } - Altera_Forum
Honored Contributor
--- Quote Start --- @BillA I'm also sending (hardware generated) UDP packets to a PC. The best results are achieved if I have a point to point connection from hardware to PC with a separate network adapter just for the hardware. For some Intel NICs (Pro 1000 series) you can increase the number of receive descriptors (see attached jpg). This reduces in my case significant loss of packets. Jens --- Quote End --- Thanks Jens - this is very helpful. My 2nd NIC is an Intel-based device and also does have this RX (and TX) buffers setting. I will do some testing. I have a reliable UDP stream protocol and it should handle lots of data without loss but I was having to stall it at (go figure) every 256 packets (which is the Intel default and might be the Broadcom default which is my other NIC). I might still have to stall at this boundary but the stalls can be farther apart which is an improvement. Bill - Altera_Forum
Honored Contributor
--- Quote Start --- Could you try to execute a small programm from any other RAM than SDRAM to check if tc_code onchip RAM (SSRAM, ...) is working correct? Then you must change the system lib properties. --- Quote End --- I definitely think I have some problem with loader or with configuration of jtag debugger or inside the project. I followed your advice and compiled your same hello_world sample. This what I obtained: case 1 Conditions: Same memory mapping as my original project; all memory sections mapped to sdram and attribute directive used to map the 2 functions to tc_code Result: same as my original project; tc_code not loaded I verified that tc_code can be writted and read case 2 Conditions: mapped .text section to sram; others to sdram Result: sram loaded and executing; functions mapped to tc_code with attribute don't. Added a (supposed) initialized variable with attribute(... tc_data): this is not initialized, too. case 3 Conditions: same as case 2 but mapped stack section to tc_data Result: same as before for code, but now the tc_data variable is correctly initialized. Now, during the loading process, I can see the tc_data addresses in the download progress log shown in ide console; before I didn't and I saw only sdram addresses. Conclusion: Apparently only section which are explicitly used in sys library properties are actually loaded. Other sections mapped with attribute directive are ignored, unless the referred section is used for anything else. Please, any help would be greatly appreciated because I've been stuck on this weird point for about a week.