Merge h. and v. passes in 4-tap SBD Neon DotProd 2D convolution
The current SBD Neon DotProd approach to 2D convolution is:
1) Filter horizontally, storing to an intermediate buffer.
2) Filter vertically and store the final output.
This patch merges the two phases for 4-tap standard bitdepth 2D
convolution to avoid storing to and re-loading from the intermediate
buffer - giving a 10-25% speedup depending on block size. Merging the
passes for 8-tap filters does not have the same benefit, so keep the
existing implementation.
Change-Id: Ic6008836d1a499ee2cd957b9db194fca5671ccb4
1 file changed