Optimize vpx_satd_neon Optimize Neon implementation of vpx_satd by using ABD and UADALP instead of ABAL and ABAL2, splitting the accumulator and using a dedicated helper function to perform the final reduction. Change-Id: Idcfa49e001b68b1dcd87c13fd9acc317a208cd2a