Methodology
Content is annotated at the item level (article, post, video) along seven independent layers — faceted annotation instead of a single flat label. The verifiable (veracity) is separated from the inferred (intent). Every taxonomy value has a citable academic source; this page is generated directly from the canonical JSON Schema (v0.2) and codebook (v0.3), so it cannot drift away from the research canon.
The canonical wording of the schema and codebook is Slovak; the English descriptions and value definitions below are translations kept in sync with the canon by an automated drift guard.
L0 — Provenance and metadata
Filled in by the collection pipeline, not by the annotator.
| field | type | values |
|---|---|---|
| item_id | text | free text / structured |
| url | text | free text / structured |
| archive_url | text | free text / structured |
| platform | single choice |
|
| media_type | single choice |
|
| source_id pseudonymised (hashed) for ordinary users | text | free text / structured |
| published_at | text | free text / structured |
| collected_at | text | free text / structured |
| language_variety | single choice |
|
| reach_snapshot | object | free text / structured |
L1 — Veracity / type of information disorder
Wardle & Derakhshan (2017): Information Disorder. Council of Europe.
| field | type | values |
|---|---|---|
| disorder_type disinformation requires falsity (L1) AND an indicator of intent/harm (L5/L6); without it, misinformation. | single choice |
|
| veracity If a claim cannot be verified from the text alone, use 'unverifiable' with low confidence. Never infer falsity from disagreement. | single choice |
|
| content_type Wardle's 7 types + not_applicable (ADR-028: the counterpart of disorder_type=not_applicable). | single choice |
|
| check_worthy Does it contain a verifiable factual claim? (filters out pure opinion) | yes/no | free text / structured |
L2 — Narrative and framing
EUvsDisinfo (EEAS) — 6 macro-narratives · Card et al. (2015): Media Frames Corpus — 14 frames.
| field | type | values |
|---|---|---|
| meta_narrative The six EUvsDisinfo macro-templates. Empty if none fits. | multi-label |
|
| frame Card et al. 2015 — the 14 MFC frames. | multi-label |
|
| primary_topic Slovak-localised seed set + emergent codes (other). | multi-label |
|
| sub_narrative | text | free text / structured |
| narrative_summary One sentence in the annotator's own words. | text | free text / structured |
| euvsdisinfo_xwalk | text | free text / structured |
L3 — Manipulation techniques
SemEval-2023 Task 3 (Piskorski et al., 2023) — 23 persuasion techniques · Da San Martino et al. (2019).
| field | type | values |
|---|---|---|
| techniques Empty if no technique is present. | multi-label |
|
L4 — Actor and behaviour
DISARM Framework (T-codes).
| field | type | values |
|---|---|---|
| source_type ADR-028: added mainstream_media (a legitimate/mainstream outlet) and commentator_influencer (an individual commentator/influencer with no political affiliation). Rejected: alternative_media (an evaluative, not a structural category). | single choice |
|
| origin | single choice |
|
| coordination_signals | single choice |
|
| disarm_ttp DISARM T-codes (actor/environment behaviour). Only when behaviour is indicated. | list | free text / structured |
| linked_entities Pseudonymised for ordinary users. | list | free text / structured |
L5 — Target and effect
Nimmo (2015): 4D — dismiss, distort, distract, dismay · DISARM.
| field | type | values |
|---|---|---|
| target_entity | single choice |
|
| intended_effect Nimmo 4D = DISARM T0075–T0078, + not_applicable (ADR-028). NEVER null: null cannot distinguish 'does not apply' from 'the model did not know'. | single choice |
|
| communicative_function | multi-label |
|
L6 — Intent (inferred)
Inferred layer — intent; annotated with low confidence, never the sole basis for classification.
| field | type | values |
|---|---|---|
| assessed_intent_to_harm | single choice |
|
| intent_confidence | single choice |
|
| intent_rationale | text | free text / structured |
L7 — Annotation process
Confidence, adjudication, audit trail.
| field | type | values |
|---|---|---|
| annotator_id e.g. 'A1', 'A2' or 'model:claude-...' | text | free text / structured |
| annotated_at | text | free text / structured |
| field_confidence Map of field -> confidence for the key fields (at least L1.veracity, L3, L5, L6). | object | free text / structured |
| notes | text | free text / structured |
| adjudication_status | single choice |
|
| gold_label | yes/no | free text / structured |
Reliability measurement
The gold corpus is built by independent annotation (coders cannot see each other's answers; a submitted annotation is immutable). Reliability is measured with Krippendorff's α per layer and per label; inter-annotator agreement forms the empirical ceiling against which the AI classifier is evaluated — not an arbitrary fixed percentage.