This research introduces the first public multi-modal dataset of 100 aligned audio-transcript pairs for Turkish scam and benign calls. It evaluates seven large language models under raw audio, automatic, and human-corrected transcript inputs, finding that transcript-based inputs outperform direct audio processing, with human correction having minimal impact.
Poster: Exploring Audio-Based Scam Detection in Turkish
from English