Forum Discussion
Altera_Forum
Honored Contributor
15 years agoIOWR and IORD with PIOs
I have a basic question concerning PIOs. Currently, I am sending 4 8-bit integers over 4 ouput PIOs from NIOS to hardware to do a calculation on them and return another 8-bit integer value. This is achieved as below.
alt_u8 val_1 = 10;
alt_u8 val_2 = 15;
alt_u8 val_3 = 20;
alt_u8 val_4 = 25;
alt_u8 result_val;
IOWR_ALTERA_AVALON_PIO_DATA(DATA_OUT1_BASE , val_1);
.
.
IOWR_ALTERA_AVALON_PIO_DATA(DATA_OUT4_BASE , val_4);
result_val = IORD_ALTERA_AVALON_PIO_DATA(RESULT_IN0_BASE,0);
Now instead of using 4 PIOs of width 8, I want to use 1 PIO of width 32 to send my four integer values. What are the IOWR commands that I should write assuming my new 32-bit PIO base address is DATA_OUT_32_BASE? Also on the hardware side where the calculation is done, the input port will now be a 32-bit input, i.e. [31:0] data_32_in;. How do I separate that input signal into its 4 component integers to use in the calculation?15 Replies
- Altera_Forum
Honored Contributor
--- Quote Start --- -When I have two IOWRs sequentially as in my code above, do these two 'writes' happen at the same time, i.e. does the hardware module who is waiting fot the data get them all at once ? --- Quote End --- In such cases the writes happen sequentially, but the time between them could be umpredictable. It depends from your hw and sw. For example you can have another Avalon master acquiring the control of the bus between them; or your program have to serve an irq. I don't think in any case the PIO would switch all at once; you need a latch module if you want to sinchronize them. --- Quote Start --- - To increase calculation speed, my next step will be to instantiate 2 calculation modules in my top level .v file, and modify the for-loops such that my image is divided into 2 (i.e. 4x5 and 4x5), and data from each sub-image goes to its respective hardware module instantiation. To do this I will need another set of PIOs (2x32bits out and 1x8bits in). Is this the right way to do it? If keep dividing my big image into more sub-images e.g 4, do I add another 2 sets of PIOs? I will end up having many sets of PIOs if my image is bigger and I want to sub-divide more! --- Quote End --- I don't understand clearly what you mean and if it is really correct for your purpose. Probably yes. And normally dividing the main task and replicating hardware in such way brings to a great improvement in speed, but also to a great increase of fpga resource utilization; you must seek the optimal configuration for your actual system. --- Quote Start --- Is there a better way to transfer the data from NIOS to the different hardware modules? Another forum member has suggested to create a wrapper module for all my instantiations and create an Avalon MM to interface to the array, but I am completely lost! Can somebody please guide me a bit through this? --- Quote End --- What the other member suggested is generally exact. PIOs are very inefficient if you want to write an external register and aMM slave interface is more convenient. But in your case I see you don't have a 'real' memory interface, being your external module totally asynchronous (and maybe combinatorial?): so I don't see any improvement in switching from PIO to Avalon MM. However giving you informatio about Avalon MM here is not feasible, since it would require a lot of space. You'll learn more by browsing Nios/sopc documentation or searching in the forum. Regards Cris - Altera_Forum
Honored Contributor
OK I understand the IOWR inefficiency, and thanks for the new code. I don't think I would have been able to do that on my own so quickly... pointers and me tend not to go well together.
Now suppose I have a 2-D array of alt_u8 of size 4 x 10, which represents an image. Then I want to apply a convolution filter to it by 'sliding' a 3x3 window on it to collect 8 values each time to send to hardware to do its calculation. Is there a smart way to use pointers here also as data are already packed in memory in the correct order with this also. I could not apply your suggested way and ended up doing it your first way (see code below).
I tried using the pointers method but am I right to say that I will need to re-assign 8 values to the alt_u8 array8b inside the two for-loops at each iteration loop and hence require more cpu work? I have new questions concerning IOWR. -When I have two IOWRs sequentially as in my code above, do these two 'writes' happen at the same time, i.e. does the hardware module who is waiting fot the data get them all at once ? - To increase calculation speed, my next step will be to instantiate 2 calculation modules in my top level .v file, and modify the for-loops such that my image is divided into 2 (i.e. 4x5 and 4x5), and data from each sub-image goes to its respective hardware module instantiation. To do this I will need another set of PIOs (2x32bits out and 1x8bits in). Is this the right way to do it? If keep dividing my big image into more sub-images e.g 4, do I add another 2 sets of PIOs? I will end up having many sets of PIOs if my image is bigger and I want to sub-divide more! Is there a better way to transfer the data from NIOS to the different hardware modules? Another forum member has suggested to create a wrapper module for all my instantiations and create an Avalon MM to interface to the array, but I am completely lost! Can somebody please guide me a bit through this?alt_u8 array_image; // fill in values for(Y=0; Y<(4-2); Y++) { for(X=0; X<(10-2); X++) { // hardware calculation IOWR_ALTERA_AVALON_PIO_DATA(DATAA_OUT1_32BITS_BASE,(array_image<<24)|(array_image<<16)|(array_image<<8)|array_image); IOWR_ALTERA_AVALON_PIO_DATA(DATAA_OUT2_32BITS_BASE,(array_image<<24)|(array_image<<16)|(array_image<<8)|array_image); // results back from hardware store_array= IORD_ALTERA_AVALON_PIO_DATA(RESULTA_IN_8BITS_BASE); } } - Altera_Forum
Honored Contributor
If you already have the 32bit value (or in general a data whose width matches the port width) the single write would be efficient.
But if you must manually build the 32bit word from single bytes (as in my example), you generally need cpu work, unless your bytes are already packed in memory in the correct order. Your idea is correct, but you must take care to declare the 16 byte data so that in memory it is equivalent to 4 32-bit words. Then you'll make a trick with pointers without requiring any cpu effort. For example: alt_u8 array8b[16]; alt_u32 * array32b; array32b = (alt_u32*)array8b; Now any reference to array32b[n] is equivalent to (array8b[n*4] | (array8b[n*4+1]<<8) | (array8b[n*4+2]<<16) | (array8b[n*4+3]<<24)) - Altera_Forum
Honored Contributor
--- Quote Start --- It would be something like this: IOWR_ALTERA_AVALON_PIO_DATA(DATA_OUT_32_BASE , (val_1<<24) | (val_2<<16) (val_3<<8) | val_4 ); But this is very inefficient! Still better the 4 byte writes. --- Quote End --- Thank you very much for the code. Could you please tell me why it is inefficient? I wanted to do it this way because I could potentially need to send 16 alt_u8 integers at a time for another calculation. Instead of having 16 8-bit PIOs, I wanted 4 32-bit ones just to make things less cumbersome in SOPC builder. So do you think I should keep my 16 separate 8-bit PIOs or is there a better way to send these integers? - Altera_Forum
Honored Contributor
It would be something like this:
IOWR_ALTERA_AVALON_PIO_DATA(DATA_OUT_32_BASE , (val_1<<24) | (val_2<<16) (val_3<<8) | val_4 ); But this is very inefficient! Still better the 4 byte writes. Regarding the hw side, every 8bit component will simply get the 8 bits it needs i.e. (Verilog) assign data_8_in1[7:0] = data_32_in[7:0]; assign data_8_in2[7:0] = data_32_in[15:8]; assign data_8_in3[7:0] = data_32_in[23:16]; assign data_8_in4[7:0] = data_32_in[31:24];