This model is from the paper "AuViRe: Audio-visual Speech Representation Reconstruction for Deepfake Temporal Localization" available at
https://huggingface.co/papers/2511.18993
This model has been pushed to the Hub using the
PytorchModelHubMixin integration: