Recovering the Longue Durée of Dissent: A Robust Framework for Cross-Temporal and Cross-Source Event Classification
- Programa:
- Sesión 7, Sesión 7
Día: viernes, 11 de septiembre de 2026
Hora: 09:00 a 10:45
Lugar: 23
This paper presents a robust methodological framework for longitudinal Protest Event Analysis (PEA) that bridges the gap between contemporary manual coding and century-long media archives. Utilizing 120 years of ABC Spain (1905–2025), we introduce a hybrid pipeline designed to overcome diachronic linguistic drift and ideological framing biases. The approach utilizes Large Language Models (LLMs) to distill a "silver standard" dataset by extracting verbatim evidence from hand-coded modern records (DISDEM-ECOPOL 2000–2020). This distilled data is used to train a Sentence Transformer Fine-Tuning (SetFit) classifier, optimized for high-precision, few-shot classification of protest features such as demands, repertoires, and organizers. By implementing an archive-specific domain adaptation step and a spatio-temporal clustering algorithm for event de-duplication, the framework ensures consistency across diverse political regimes and editorial shifts. Validation against contemporary gold standards and historical samples demonstrates that this pipeline effectively recovers the longue durée of social contention with unprecedented semantic granularity.
Palabras clave: Protest Event Analysis (PEA), Silver Standard Distillation, SetFit (Sentence Transformer Fine-Tuning), Computational Social Science, Digital Humanities