The recovery of latent sparse spike trains from multichannel observations is a convolutive blind source separation problem central to many applications in neural signal processing, where the goal is to decompose recordings into individual neural firing patterns. Existing approaches rely on time-domain signal extension followed by independent component analysis, providing point estimates but no uncertainty quantification. We propose a hierarchical Bayesian framework formulated directly in convolutional space. Each latent source is modelled as a Bernoulli-Gaussian sparse representation convolved with a channel-specific finite impulse response, with Student-t residuals for robustness to heavytailed noise. Inference proceeds via variational EM: a mean-field E-step yields closed-form posteriors over spike presence and amplitude, while the M-step solves a ridge-regularised deconvolution problem in the frequency domain via preconditioned conjugate gradient. We evaluate the method on simulated and experimental high-density surface electromyography data, demonstrating spike train estimates with negligible baseline noise compared to state-of-the-art methods and reduced memory requirements.