[
https://issues.apache.org/jira/browse/SINGA-80?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14960832#comment-14960832
]
wangwei commented on SINGA-80:
------------------------------
When and where to indicate the running device, i.e., cpu or gpu, should be
considered.
One approach is making Layer a class template with argument xpu = cpu or gpu.
If the blob level math function is also declared as function templates, e.g.,
{code}
template<xpu>
void Add(const Blob& A, const Blob& B, Blob* C);
{code}
the layer implementation would be simple (in dependent of the running devices),
e.g.,
{code}
template<xpu>
void Layer::ComputeFeature() {
Add<xpu>(A, B, &C);
}
{code}
For some layers, e.g., input layers, they can only run on CPU, then we fix the
xpu to cpu in the implementation.
{code}
template<xpu>
void Layer::ComputeFeature() {
Add<cpu>(A, B, &C);
}
{code}
Another problem to consider is address level functions. Do we use function
templates?
If the implementation of blob level functions are *exactly the same except the
calling of address level functions*, then we can use function templates for
address level functions and calling them in the Blob math functions, e.g.,
{code}
template<xpu>
void Add(const Blob& A, const Blob& B, Blob* C) {
add<xpu>(A.dptr, B.dptr, C->dptr);
}
{code}
Otherwise, we need to differentiate the implementations by
{code}
template<>
void Add<cpu>(const Blob& A, const Blob& B, Blob* C) {
...// operations specific for cpu
cpu_add(A.dptr, B.dptr, C->dptr);
}
template<>
void Add<gpu>(const Blob& A, const Blob& B, Blob* C) {
...// operations specific for gpu
gpu_add(A.dptr, B.dptr, C->dptr);
}
{code}
> New Blob Level and Address Level Math Operation Interface
> ---------------------------------------------------------
>
> Key: SINGA-80
> URL: https://issues.apache.org/jira/browse/SINGA-80
> Project: Singa
> Issue Type: Improvement
> Reporter: Jinyang Gao
> Assignee: Jinyang Gao
> Original Estimate: 672h
> Remaining Estimate: 672h
>
> We are going to provide a new two level math interface to replace the current
> mshadow. The higher blob level interface is going to be used by layer level.
> It is xpu transparent, and will support general matrix, element-wise,
> reduce/expand, pack/unpack operations and etc. in blob level. There is no
> further need to transfer the blob object into tensor object before math
> operation. The lower address level interface is going to support efficient
> cpu/gpu computing task on simple data array.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)