Methodology

Content is annotated at the item level (article, post, video) along seven independent layers — faceted annotation instead of a single flat label. The verifiable (veracity) is separated from the inferred (intent). Every taxonomy value has a citable academic source; this page is generated directly from the canonical JSON Schema (v0.2) and codebook (v0.3), so it cannot drift away from the research canon.

The canonical wording of the schema and codebook is Slovak; the English descriptions and value definitions below are translations kept in sync with the canon by an automated drift guard.

L0 — Provenance and metadata

Filled in by the collection pipeline, not by the annotator.

fieldtypevalues
item_idtextfree text / structured
urltextfree text / structured
archive_urltextfree text / structured
platformsingle choice
  • facebook
  • x
  • telegram
  • youtube
  • web
  • ine
media_typesingle choice
  • text
  • image
  • video
  • audio
  • multimodal
source_id

pseudonymised (hashed) for ordinary users

textfree text / structured
published_attextfree text / structured
collected_attextfree text / structured
language_varietysingle choice
  • sk
  • cz
  • sk-dialect
  • mixed
  • ine
reach_snapshotobjectfree text / structured

L1 — Veracity / type of information disorder

Wardle & Derakhshan (2017): Information Disorder. Council of Europe.

fieldtypevalues
disorder_type

disinformation requires falsity (L1) AND an indicator of intent/harm (L5/L6); without it, misinformation.

single choice
  • disinformation
  • misinformation
  • malinformation
  • not_applicable
veracity

If a claim cannot be verified from the text alone, use 'unverifiable' with low confidence. Never infer falsity from disagreement.

single choice
  • true
  • mostly_true
  • misleading
  • false
  • fabricated
  • unverifiable
  • opinion_not_factual
content_type

Wardle's 7 types + not_applicable (ADR-028: the counterpart of disorder_type=not_applicable).

single choice
  • satire_parody
  • false_connection
  • misleading_content
  • false_context
  • imposter_content
  • manipulated_content
  • fabricated_content
  • not_applicable
check_worthy

Does it contain a verifiable factual claim? (filters out pure opinion)

yes/nofree text / structured

L2 — Narrative and framing

EUvsDisinfo (EEAS) — 6 macro-narratives · Card et al. (2015): Media Frames Corpus — 14 frames.

fieldtypevalues
meta_narrative

The six EUvsDisinfo macro-templates. Empty if none fits.

multi-label
  • elites_vs_people — Elites vs. the people — corrupt “elites” detached from “ordinary people”; secret elites; electoral fraud
  • threatened_values — Threatened values — the decadent West; traditional values under attack from liberalism/LGBTI/feminism; Russia as the guardian of decency
  • lost_sovereignty — Lost sovereignty — countries are no longer “truly sovereign”; ruled by the USA/Brussels; the “colour revolutions” trope
  • imminent_collapse — Imminent collapse — the West/EU/Ukraine will soon collapse economically/politically/morally
  • hahaganda — Hahaganda — ridicule used to discredit both the target and democracy itself
  • nazism_accusations — Nazism accusations — labelling opponents (especially Ukraine) as “Nazis/fascists”
frame

Card et al. 2015 — the 14 MFC frames.

multi-label
  • Economic
  • Capacity_and_resources
  • Morality
  • Fairness_and_equality
  • Legality_Constitutionality
  • Policy_prescription
  • Crime_and_punishment
  • Security_and_defense
  • Health_and_safety
  • Quality_of_life
  • Cultural_identity
  • Public_opinion
  • Political
  • External_regulation_and_reputation
primary_topic

Slovak-localised seed set + emergent codes (other).

multi-label
  • war_ukraine — War in Ukraine — the West/Ukraine is to blame; “the war began in 2014 in Donbas”; Ukraine = a fascist state; “Russophobia” — nazism_accusations, lost_sovereignty
  • anti_nato — Against NATO — NATO is aggressive and drags us into war; a US instrument; enlargement = betrayal of Russia — lost_sovereignty, imminent_collapse
  • anti_eu — Against the EU — the “Brussels diktat”; nonsensical regulations; the EU run from Washington — lost_sovereignty, elites_vs_people
  • sanctions_energy_economy — Sanctions / energy — sanctions hurt Europe more than Russia; expensive energy — imminent_collapse
  • sovereignty_foreign_policy — Sovereignty / foreign policy — integration made Slovakia a “servant of foreign interests” — lost_sovereignty
  • traditional_values — Traditional values — Russia as the “big brother” and protector of tradition — threatened_values
  • gender_lgbti — Gender / LGBTI — “gender ideology”; a threat to the family — threatened_values
  • migration — Migration — migrants as a threat to identity and values — threatened_values
  • health_antivax_covid — Health / antivax / COVID — the “infodemic”; bio-laboratories; vaccine conspiracies — elites_vs_people
  • elections_delegitimization — Delegitimisation of elections — electoral fraud; questioning the results — elites_vs_people, hahaganda
  • distrust_institutions — Distrust of institutions — distrust of the state/police (e.g. after the 2024 assassination attempt on Fico) — elites_vs_people
  • anti_media — Against the media — scepticism towards independent media; “the mainstream lies” — elites_vs_people, hahaganda
  • anti_americanism — Anti-Americanism — the USA as the real hegemon/aggressor — lost_sovereignty
  • globalist_elite_conspiracy — Globalist conspiracy — Soros; the “Great Reset”; secret elites — elites_vs_people
  • other
sub_narrativetextfree text / structured
narrative_summary

One sentence in the annotator's own words.

textfree text / structured
euvsdisinfo_xwalktextfree text / structured

L3 — Manipulation techniques

SemEval-2023 Task 3 (Piskorski et al., 2023) — 23 persuasion techniques · Da San Martino et al. (2019).

fieldtypevalues
techniques

Empty if no technique is present.

multi-label
  • Name_Calling-Labeling — Name Calling / Labeling — labelling the object with a term the audience hates or adores — Name-calling, Labeling, Demonizing the enemy, Stereotyping
  • Doubt — Doubt — questioning the credibility of a person/institution without evidence — (part of FUD)
  • Guilt_by_Association — Guilt by Association — discrediting through association with a hated person/group — Guilt by association
  • Appeal_to_Hypocrisy — Appeal to Hypocrisy — “you did it too” — attacking the opponent's consistency — (absent in DIMARL)
  • Questioning_the_Reputation — Questioning the Reputation — attacking reputation/credentials rather than the argument — Smears, Ad hominem, False accusations
  • Appeal_to_Authority — Appeal to Authority — the claim is true because an authority says so — Appeal to authority, Testimonial, Third party technique
  • Appeal_to_Popularity — Appeal to Popularity — “everyone does it” / the majority supports it — Bandwagon
  • Appeal_to_Values — Appeal to Values — an appeal to shared values (freedom, homeland…) — Glittering generalities, Virtue words
  • Appeal_to_Fear-Prejudice — Appeal to Fear / Prejudice — evoking fear/prejudice instead of an argument — Appeal to fear, Appeal to prejudice
  • Flag_Waving — Flag Waving — an appeal to patriotism / group identity — Flag-waving
  • Straw_Man — Straw Man — distorting the opponent's position into an easier target — Straw man
  • Red_Herring — Red Herring — diverting attention to an irrelevant topic — Red herring
  • Whataboutism — Whataboutism — a counter-accusation instead of an answer — Whataboutism
  • Causal_Oversimplification — Causal Oversimplification — a single cause for a complex phenomenon — Oversimplification
  • False_Dilemma-No_Choice — False Dilemma / No Choice — only two options, as if no others existed — Black-and-white fallacy
  • Consequential_Oversimplification — Consequential Oversimplification — “if A, the chain B, C, D inevitably follows” (slippery slope) — (absent in DIMARL)
  • Slogans — Slogans — a short, punchy catchphrase instead of an argument — Slogans
  • Appeal_to_Time — Appeal to Time — “now or never” — urgency pressure — (absent in DIMARL)
  • Conversation_Killer — Conversation Killer — a phrase that shuts down the discussion — Thought-terminating cliché
  • Loaded_Language — Loaded Language — emotionally charged words — Loaded language, Dysphemism, Euphemism
  • Repetition — Repetition — repeating the same message over and over — Repetition / Ad nauseam
  • Exaggeration-Minimisation — Exaggeration / Minimisation — inflating or downplaying — Exaggeration, Minimisation
  • Obfuscation-Vagueness-Confusion — Obfuscation – Vagueness / Confusion — deliberate vagueness/confusion — Obfuscation, Intentional vagueness
  • Scapegoating

L4 — Actor and behaviour

DISARM Framework (T-codes).

fieldtypevalues
source_type

ADR-028: added mainstream_media (a legitimate/mainstream outlet) and commentator_influencer (an individual commentator/influencer with no political affiliation). Rejected: alternative_media (an evaluative, not a structural category).

single choice
  • state_affiliated
  • known_disinfo_outlet
  • political_actor
  • influencer
  • anonymous
  • suspected_inauthentic
  • ordinary_user
  • mainstream_media
  • commentator_influencer
  • unknown
originsingle choice
  • domestic
  • foreign
  • unknown
coordination_signalssingle choice
  • yes
  • no
  • unknown
disarm_ttp

DISARM T-codes (actor/environment behaviour). Only when behaviour is indicated.

listfree text / structured
linked_entities

Pseudonymised for ordinary users.

listfree text / structured

L5 — Target and effect

Nimmo (2015): 4D — dismiss, distort, distract, dismay · DISARM.

fieldtypevalues
target_entitysingle choice
  • institution
  • EU
  • NATO
  • ukraine
  • government
  • specific_person
  • science_experts
  • minority_group
  • election
  • other
intended_effect

Nimmo 4D = DISARM T0075–T0078, + not_applicable (ADR-028). NEVER null: null cannot distinguish 'does not apply' from 'the model did not know'.

single choice
  • dismiss
  • distort
  • distract
  • dismay
  • not_applicable
communicative_functionmulti-label
  • discredit
  • mobilize
  • demobilize
  • monetize
  • polarize
  • emotional_arousal

L6 — Intent (inferred)

Inferred layer — intent; annotated with low confidence, never the sole basis for classification.

fieldtypevalues
assessed_intent_to_harmsingle choice
  • yes
  • no
  • unknown
intent_confidencesingle choice
  • low
  • medium
  • high
intent_rationaletextfree text / structured

L7 — Annotation process

Confidence, adjudication, audit trail.

fieldtypevalues
annotator_id

e.g. 'A1', 'A2' or 'model:claude-...'

textfree text / structured
annotated_attextfree text / structured
field_confidence

Map of field -> confidence for the key fields (at least L1.veracity, L3, L5, L6).

objectfree text / structured
notestextfree text / structured
adjudication_statussingle choice
  • pending
  • resolved
gold_labelyes/nofree text / structured

Reliability measurement

The gold corpus is built by independent annotation (coders cannot see each other's answers; a submitted annotation is immutable). Reliability is measured with Krippendorff's α per layer and per label; inter-annotator agreement forms the empirical ceiling against which the AI classifier is evaluated — not an arbitrary fixed percentage.