Replace ReduceDimsOps math::Gemv with CUDA reduction kernel. 5.6x speed up. Summary: This reduces runtime from 1.54757 ms/iter -> 0.273687 ms/iter for an 100 parallel reductions each of size 100000. Reviewed By: akyrola Differential Revision: D5471324 fbshipit-source-id: 626cabb8249fb4655275648fae2738cb739e1a72
Caffe2 is a lightweight, modular, and scalable deep learning framework. Building on the original Caffe, Caffe2 is designed with expression, speed, and modularity in mind.
Caffe2 research award competition request for proposals
Please use Github issues (https://github.com/caffe2/caffe2/issues) to ask questions, report bugs, and request new features.
Please participate in our survey (https://www.surveymonkey.com/r/caffe2). We will send you information about new releases and special developer events/webinars.
Caffe2 is released under the BSD 2-Clause license.