Forum Discussion
Code in tightly coupled memory
Hi all,
I placed a function in a section different from the main code. More exactly, my code is mapped to sdram and I linked a function testcall() into a tightly coupled memory section which I previoulsy defined in sopc builder. So I declared: int testcall(int n) __attribute__ ((section (".tc_code"))); Then, following the suggestion I found in another thread I called ALT_LOAD_SECTION_BY_NAME(tc_code) before calling testcall(). My questions are: - is the ALT_LOAD_SECTION_BY_NAME(tc_code) mandatory? The code seems to work even if I don't use it. - if tc_code section is located in internal fpga memory (cyclone III M9k blocks), how much is the code speed improvement I could roughly obtain with respect to sdram (assume I'm using 2 tse with sgdma). Thank you Cris37 Replies
- Altera_Forum
Honored Contributor
I guess that ALT_LOAD_SECTION_BY_NAME(tc_code) copies the code from where ever the normal loader put it into the M9K block.
Whether this is necessary depends on how the program image was linked, and how it was loaded. The JTAG loader reads the elf program headers - so can write into multiple disjoint areas - in which case no copy is required (if the program headers are 'correct'). In other cases, a more simple loader might have written the code somewhere else - so the copy is needed. Putting code (and data) in tightly coupled memory areas gives the same access times as if the data were resident in the instruction/data cache. - Altera_Forum
Honored Contributor
Thank you for your answer and the explanation of the ALT_LOAD_SECTION behavior.
--- Quote Start --- Putting code (and data) in tightly coupled memory areas gives the same access times as if the data were resident in the instruction/data cache. --- Quote End --- This was already clear to me. I simply wondered if I can expect any significative speed improvement in placing frequently accessed code/data in a dedicated tightly coupled memory. Note that I already use a Nios II/f code with cache: anyway I want very fast execution of this function, so I'd like to exclude cache delays due to loading/flushing. I add one more issue to this thread: I tried to place in tc_code section a more complex function, namely the actual function I need to speed up. Now the processor gets stuck whenever this funciton is called! The previous call to testcall() is still present and it has no problem. Why? Regards Cris - Altera_Forum
Honored Contributor
If you want a function to execute as fast as possible you probably need to inspect the generated code and examine the control flow and memory accesses to avoid execution stalls and unwanted memory accesses.
Depending on the code I'd guess you can gain 20-30% by optimising the C so that gcc generates better code. - Altera_Forum
Honored Contributor
I want give just my experience on working with tightly coupled memory. I'm using TCM for an ISR and a few time critical functions. First I forgot to put the variables used by the functions in TCM also outside from SDRAM to TCM ( __attribute__ directive just as for the variables).
In my case it was necessary to disable IRQs during functio execution:
Apart from performance increase another benefit of transfering functions in TCM is an uninterrupted access to SDRAM for other (custom) bus masters. Jensalt_irq_context context __attribute__ ((section (".onchip_ram"))); context = alt_irq_disable_all(); //time critical code alt_irq_enable_all(context); - Altera_Forum
Honored Contributor
hi Jens,
In your case, was disabling interrupts for TCM code necessary for the correct execution of code itself or only for performance? In my case I'd allow interrupts since I need to speed up the "average" function execution time and I don't mind if it is sometimes interrupted; think that this function is almost always running, so it DOES need to be interrupted. Can this be the reason why the simple function works while the complete one doesn't? Same question above about placing variables in TCM. I finally agree with your last remark: I'd expect performance increase not for the TCM itself but because sdram is now intensively accessed by other masters through dma. Cris - Altera_Forum
Honored Contributor
Cris, I use the function in TCM to cyclical control and setup SGDMA descriptors. Data from a sensor must be copied in limited time to SDRAM. If I allow IRQs then I don't receive all the data from sensor. In this time any other master access on the SDRAM leads to loss of data.
Furthermore I use global variables for that function because stack and heap are in SDRAM. I mean functionality should not affected if IRQs allowed or not but I couldn't test it. An other effect I had in conjunction with dual port RAM. First I put TCM and descriptors together in one dual port RAM. In this case I also had loss of data. How do you create the tc_code section? The TCM section I'm using has the same name like the SOPC component (onchip_ram). This naming has effects to the generated linker script. If you will use additional sections than you have to provide a custom linker script. Jens - Altera_Forum
Honored Contributor
--- Quote Start --- Data from a sensor must be copied in limited time to SDRAM. If I allow IRQs then I don't receive all the data from sensor. In this time any other master access on the SDRAM leads to loss of data. --- Quote End --- So in your case the IRQ issue would be present even if you didn't use TCM and mapped the code into sdram, sram or anything else. Right? I tried with irq disable and infact this is not my case. --- Quote Start --- How do you create the tc_code section? The TCM section I'm using has the same name like the SOPC component (onchip_ram). --- Quote End --- That's what I did. I still have the same problem: simple function in TCM works, complete function hangs. - simple function: for loop which sums the first n integers, where n is the function parameter; 0x2C bytes code - complete function: 0xbc bytes code; this function works perfectly if I map it into sdram - tc_code : 8kbytes space I will try now to progressively increase the size of the simple test function in order to find out if the problem is with code size or with function content Cris - Altera_Forum
Honored Contributor
Yes, thats right. In my case it was a timing problem. Send me your functions if you want. I would have a look on these.
Jens - Altera_Forum
Honored Contributor
This works:
I inserted the IOWR and printf to increase code size and have a call to a function in sdram, like the 'real' case. This DOES NOT work:int testcall(int n) __attribute__ ((section (".tc_code"))); int testcall(int n) { int i; i = n; while (i > 0) { n += i--; IOWR_ALTERA_AVALON_PIO_DATA(IO24V_PIO_BASE, n); printf("."); } IOWR_ALTERA_AVALON_PIO_DATA(IO24V_PIO_BASE, 0x55); return n; }
Remarks: - buffer_tx resides in sdram - pkt_send() works perfectly when I remove the attribute directive - testcall() is called in the very beginning of program; pkt_send after uC RTOS tasks have been initialized Thank you for any help Crisint pkt_send(char *data, int len) __attribute__ ((section (".tc_code"))); struct buffer_t { int len; unsigned char data; } buffer_tx; int pkt_send(char *data, int len) { short index; // is next buffer valid? index = pkt_send_count & 0x07; if (buffer_tx.len != 0) return 0; if (len >= 0x600) len = 0x600; memcpy(buffer_tx.data, data, len); buffer_tx.len = len; ++pkt_send_count; return len; } - Altera_Forum
Honored Contributor
More info:
I switched to Debug mode and neither the testcall() function works anymore; nor I can debug inside the functions located in tcm section (but maybe this is a normal debugger limit)