Hi chris, thank you for your response.
Yes. I agree with you basically. According to my current observation,
too much overhead on the communication between DSP and ARM involves two
items:
(1) too many time consumed on cache management.
I have tried ever to "trick" the framework that some input / output
buffers is very smaller than their actual sizes.
The average overhead of ARM/DSP communication can be reduced from
0.9 => 0.3 ms.
(2) the priority of ARM application.
in my application, i didn't introduce any additional
process/thread. it is just a very simple single-process application ---
in the main() function, it do almost everything.
Thus, the priority of this process will be relatively low. I have
tried to increase the application prority by calling setpriority()
function. it seems that jumping points with very large overhead will
occurs less frequently. and at these points, the overhead tend to be
smaller, e.g., generally, it will be 2ms overhead or less less
frequently, it will be 4ms overhead.
More information about my application:
1. when running my application, i didn't run any other task. I just
start the board, then log in, then run my application. All these things
are doned via a serial-port console window. (so, there are no other user
thread running, except the kernel thread).
2. the size of my input/output buffers are always in a relatively
constant order of magnitude. they will not become very large or very
small.
Any further sugestion? And can you give me a more detail explanation
about your comments:
"This could be because an ARM-side driver is holding off interrupts for
a "long time" (i.e. keeping the DSP->ARM interrupt from being
serviced)... or even simply that a higher priority thread is running
ahead of your decode application. Anything else going on in the
system?"
How to prevent "ARM-side driver is holding off interrupts"?
i have tried increase the priorty of my decode applcation, and i believe
that there are no other user thread running, except the kernel thread.
In this situation, how can i do furtherly?
-----Original Message-----
From: Ring, Chris [mailto:[EMAIL PROTECTED]
Sent: Tuesday, May 08, 2007 4:10 AM
To: X. Zhou; [email protected]
Cc: Huo Yan Jenny
Subject: RE: the communication between DSP and ARM takes up too much
time!
1 ms overhead is on the expected order of magnitude we've
observed as well for video codecs. This is dominated by the cache
maintenance of getting the very large video buffers from the ARM to the
DSP and back. The actual "processor context switch" overhead (ARM->DSP
and back) is typically less than 200 microseconds. But there's
significant overhead making the memory "look right" when taking the DSP
cache into account.
However, this doesn't explain why you're seeing up to 7 ms
overhead at times (unless your buffers are occasionally _very_, _very_
large). This could be because an ARM-side driver is holding off
interrupts for a "long time" (i.e. keeping the DSP->ARM interrupt from
being serviced)... or even simply that a higher priority thread is
running ahead of your decode application. Anything else going on in the
system?
[ Note that in many codec implementations, this cache
maintenance is unnecessary (e.g. if the codec never accesses the buffers
with the CPU and instead uses DMA to access the buffers). However, to
be functionally correct _always_ the Codec Engine framework
pessimistically manages the cache "just in case" the codec used the CPU
to access it. We're taking steps in future releases of Codec Engine to
enable more efficient cache maintenance for experts who know how the
codecs behave.
And in the xDM 1.00 interfaces, we added an extra .accessMask
field to the buffer descriptors (e.g. XDM1_BufDesc) so the algorithm can
indicate _how_ it accessed the buffer, and the frameworks can then be
more efficient in managing the cache for that buffer. ]
--- Read below this line at your own risk --- ;)
The cache maintenance overhead is a factor of how big the
buffers are and how many there are. Here's a trick that I _don't_
recommend, but some customers have used successfully...
If you know your codec never accesses the data buffers with the
DSP CPU (e.g. only uses DMA), and you know the codecs don't misbehave if
the buffer sizes are incorrectly set by the app(!), you may be able to
"trick" the cache maintainers (i.e. the DSP-side VISA skeletons) into
thinking the buffers are very small so their cache maintenance overhead
is much, much less. You could set the XDM_BufDesc.bufSizes[] to very
small (e.g. 128 bytes is a single cache page), and the DSP-side
skeletons will operate much faster - since they're not managing the
cache for the entire "real" buffer.
Use this tip _only_ if you know the implementation details of
your codec! It's not a supported usage of the VISA APIs, and cache
incoherency is often really hard to debug... but this tip might help
some experts squeak out a few extra cycles. :)
Chris
________________________________
From:
[EMAIL PROTECTED]
[mailto:[EMAIL PROTECTED] On
Behalf Of X. Zhou
Sent: Sunday, May 06, 2007 9:20 PM
To: [email protected]
Cc: Huo Yan Jenny
Subject: the communication between DSP and ARM takes up
too much time!
Hi all, I encountered a problem needs your help.
I have developped a H.264 video decoder on DVEVM6446
board. It runns on DSP side and I use the codec egine mechansim to
control this decoder from ARM side.
For performance evaluation, i count the time consumed by
this decoder both from ARM side and DSP side.
On DSP side, I use the cycle hardware register to
statistic the processing time for each frame. e.g.:
H264VIDEODEC_process( )
{
.....
start = cycle();
video_dec_one_frame(); //main function
for decoding one frame
end = cycle();
decDspTime = diffcycle(start, end);
....
}
On ARM side, I use the gettimeofday() function to
statistic the processing time for each frame.
.....
start = gettimeofday();
VIDDEC_process();
end = gettimeofday();
encDspTime = difftime(start, end);
.....
It seems that in most cases, the statistical time on ARM
side is close to the statistical time on DSP side, commonly, the former
is about 1ms larger than the latter for a CIF decoder.
But in some occasional cases, the statistical time on
ARM side is very very larger than the time on DSP side. e.g., for some
frames, the former is 16ms, the latter is 9ms!!!! It is unacceptable
for my real-time requirement!
What is the cause for this case? which module take up
the additional time, i.e.: (16ms - 9ms) ?
As i know, there is a communication process between DSP
and ARM when ARM receives VIDDEC_process() function call. Is this part
takes too much time? and why and why is it so large? How can i decrease
it?
Any response is appreiated!
BR,
ZhouXiao
This message (including any attachments) is for the named addressee(s)'s
use only. It may contain
sensitive, confidential, private proprietary or legally privileged
information intended for a
specific individual and purpose, and is protected by law. If you are not
the intended recipient,
please immediately delete it and all copies of it from your system,
destroy any hard copies of it
and notify the sender. Any use, disclosure, copying, or distribution of
this message and/or any
attachments is strictly prohibited.
This message (including any attachments) is for the named addressee(s)'s use
only. It may contain
sensitive, confidential, private proprietary or legally privileged information
intended for a
specific individual and purpose, and is protected by law. If you are not the
intended recipient,
please immediately delete it and all copies of it from your system, destroy any
hard copies of it
and notify the sender. Any use, disclosure, copying, or distribution of this
message and/or any
attachments is strictly prohibited.
_______________________________________________
Davinci-linux-open-source mailing list
[email protected]
http://linux.davincidsp.com/mailman/listinfo/davinci-linux-open-source