1 ms overhead is on the expected order of magnitude we've observed as
well for video codecs.  This is dominated by the cache maintenance of
getting the very large video buffers from the ARM to the DSP and back.
The actual "processor context switch" overhead (ARM->DSP and back) is
typically less than 200 microseconds.  But there's significant overhead
making the memory "look right" when taking the DSP cache into account.
 
However, this doesn't explain why you're seeing up to 7 ms overhead at
times (unless your buffers are occasionally _very_, _very_ large).  This
could be because an ARM-side driver is holding off interrupts for a
"long time" (i.e. keeping the DSP->ARM interrupt from being serviced)...
or even simply that a higher priority thread is running ahead of your
decode application.  Anything else going on in the system?
 
[ Note that in many codec implementations, this cache maintenance is
unnecessary (e.g. if the codec never accesses the buffers with the CPU
and instead uses DMA to access the buffers).  However, to be
functionally correct _always_ the Codec Engine framework pessimistically
manages the cache "just in case" the codec used the CPU to access it.
We're taking steps in future releases of Codec Engine to enable more
efficient cache maintenance for experts who know how the codecs behave.
 
And in the xDM 1.00 interfaces, we added an extra .accessMask field to
the buffer descriptors (e.g. XDM1_BufDesc) so the algorithm can indicate
_how_ it accessed the buffer, and the frameworks can then be more
efficient in managing the cache for that buffer. ]
 
 
--- Read below this line at your own risk ---  ;)
 
The cache maintenance overhead is a factor of how big the buffers are
and how many there are.  Here's a trick that I _don't_ recommend, but
some customers have used successfully...
 
If you know your codec never accesses the data buffers with the DSP CPU
(e.g. only uses DMA), and you know the codecs don't misbehave if the
buffer sizes are incorrectly set by the app(!), you may be able to
"trick" the cache maintainers (i.e. the DSP-side VISA skeletons) into
thinking the buffers are very small so their cache maintenance overhead
is much, much less.  You could set the XDM_BufDesc.bufSizes[] to very
small (e.g. 128 bytes is a single cache page), and the DSP-side
skeletons will operate much faster - since they're not managing the
cache for the entire "real" buffer.
 
Use this tip _only_ if you know the implementation details of your
codec!  It's not a supported usage of the VISA APIs, and cache
incoherency is often really hard to debug... but this tip might help
some experts squeak out a few extra cycles.  :)
 
Chris


________________________________

        From: [EMAIL PROTECTED]
[mailto:[EMAIL PROTECTED] On
Behalf Of X. Zhou
        Sent: Sunday, May 06, 2007 9:20 PM
        To: [email protected]
        Cc: Huo Yan Jenny
        Subject: the communication between DSP and ARM takes up too much
time!
        
        
        Hi all,  I encountered a problem needs your help.
         
        I have developped a H.264 video decoder on DVEVM6446 board.   It
runns on DSP side and I use the codec egine mechansim to control this
decoder from ARM side.
         
        For performance evaluation, i count the time consumed by this
decoder both from ARM side and DSP side.
         
        On DSP side, I use the cycle hardware register to statistic the
processing time for each frame. e.g.:
                H264VIDEODEC_process( )
                {
                      .....
                     start = cycle();
                     video_dec_one_frame();    //main function for
decoding one frame
                     end = cycle();
                     decDspTime = diffcycle(start, end);
                     ....
                }
         
        On ARM side, I use the gettimeofday() function to statistic the
processing time for each frame.
               .....
               start = gettimeofday();
               VIDDEC_process();
        
               end = gettimeofday();
               encDspTime = difftime(start, end);
               .....
         
        It seems that in most cases, the statistical time on ARM side is
close to the statistical time on DSP side, commonly, the former is about
1ms larger than the latter for a CIF decoder.
         
        But in some occasional cases,  the statistical time on ARM side
is very very larger than the time on DSP side. e.g., for some frames,
the former is 16ms, the latter is 9ms!!!!  It is unacceptable for my
real-time requirement!   
        What is the cause for this case?  which module take up the
additional time, i.e.: (16ms - 9ms) ? 
        As i know,  there is a communication process between DSP and ARM
when ARM receives VIDDEC_process() function call. Is this part takes too
much time? and why and why is it so large?  How can i decrease it?  
         
         
        Any response is appreiated!
         
        BR,
        ZhouXiao
         
This message (including any attachments) is for the named addressee(s)'s
use only. It may contain
sensitive, confidential, private proprietary or legally privileged
information intended for a
specific individual and purpose, and is protected by law. If you are not
the intended recipient,
please immediately delete it and all copies of it from your system,
destroy any hard copies of it
and notify the sender. Any use, disclosure, copying, or distribution of
this message and/or any
attachments is strictly prohibited.


        

_______________________________________________
Davinci-linux-open-source mailing list
[email protected]
http://linux.davincidsp.com/mailman/listinfo/davinci-linux-open-source

Reply via email to