An undergraduate student on the Hugging Face forums discusses applying concepts from the VAPO paper to the Qwen 2.5 Omni 3B model to address hallucination issues during audio transcription with slide context.
- The VAPO paper identifies that flattening and concatenating different token types in Omni-modal Large Language Models causes visual interference, leading to hallucinated text in transcripts.
- VAPO solves this using a reinforcement learning paradigm with a 'temporally decoupled policy' that separates visual prior extraction from transcription.
- The student asks if architectural modifications to Qwen 2.5 Omni 3B, such as capturing embeddings before the thinker block or altering the thinker block itself, can prevent visual tokens from interfering with audio tokens.
The post seeks technical guidance on implementing these decoupling strategies within the existing Qwen architecture.