Skip to content
Don't Feed the TrollsGet help

How conversations turn toxic

A discussion going bad is treated as something the conversation does over time, rather than as a property of any single comment in it.

Evidence: Single study One study, or a small number by the same group. Treat as suggestive.

Underneath most bad threads is a question that sounds unanswerable: could anyone have seen this coming? One line of research treats it as a technical problem — and the way it sets the problem up turns out to be more useful to a reader than its results are.

The word means something different here

Researchers here use "derailment" for a conversation that turns toxic as it goes on. This site uses derailing for something else: steering a discussion away from its subject so the original point is never addressed. The two senses share a word and very little else, which is why the work described on this page is deliberately not cited on that one.

What the work does

Earlier efforts, as these authors describe them, detected antisocial behaviour after the fact, by analysing single comments in isolation. [1] The aim here was the opposite: to give human moderators notice before a conversation turns toxic rather than after. [2]

That aim forces a particular shape on the problem. It means modelling derailment as an emerging property of a conversation rather than as an isolated utterance-level event. [3] Because conversations are dynamic, a forecasting model has to capture the flow of the discussion rather than properties of individual comments. [4] And it has to work without knowing how long it has: real conversations have an unknown horizon, and can end or derail at any time. [5]

Applied to two new datasets of online conversations labelled for antisocial events, the model outperformed state-of-the-art systems at forecasting derailment. [6]

Why the framing matters more than the result

Outperforming state-of-the-art systems is a comparison between systems. It is not a statement of how often the model is right, and the archived page this site cites does not give one — so treat the result as evidence that the problem is tractable, not as a number you can rely on.

The framing is the part that survives. If a thread going bad is a property of the conversation rather than of a comment, then "which of these people is the troll?" is the wrong first question, and "what is this thread doing?" is a better one.

That is also where this meets the other page in this section. One author is common to both this paper and a study of what triggers trolling in the first place. [7] In that study, mood and discussion context together can explain trolling behaviour better than an individual's history of trolling. [8] Two separate pieces of work, both finding that the conversation carries more information than the person does. Why people troll covers that study in full.

What it does not give you

It gives you no cues. Nothing here names a word, a turn or a tone that means a thread is about to go wrong, and a forecast made over a labelled dataset is not advice about the argument you are in tonight. If you want the practical version, it is not a checklist: it is the same conclusion the other page reaches, that what the thread is doing to everyone reading matters more than what one account is doing. That reading is this site's own, not a finding of either paper.

What this does not show

This is one line of work, evaluated on two datasets, and what it produced is a forecasting model rather than a finding about people. Outperforming other systems is a comparison between systems, not a measure of how often it is right, and the archived page this site cites gives no accuracy figure at all. Nothing here tells a reader which words or turns to watch for, and nothing here supports a judgement about the person you are arguing with. Note also that "derailment" means something narrower here than the word does on the rest of this site.

Also in this section

  • Why people troll Two very different explanations — stable personality traits, and ordinary people in bad circumstances — and the evidence supports both.

Sources

Last checked . Each quote below is verified against an archived copy of its source every time the site is built.

  1. [1] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    recent efforts mostly focused on detecting antisocial behavior after the fact, by analyzing single comments in isolation
  2. [2] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    to provide more timely notice to human moderators, a system needs to preemptively detect that a conversation is heading towards derailment before it actually turns toxic
  3. [3] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    modeling derailment as an emerging property of a conversation rather than as an isolated utterance-level event
  4. [4] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    since conversations are dynamic, a forecasting model needs to capture the flow of the discussion, rather than properties of individual comments
  5. [5] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    real conversations have an unknown horizon: they can end or derail at any time
  6. [6] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    by applying this model to two new diverse datasets of online conversations with labels for antisocial events, we show that it outperforms state-of-the-art systems at forecasting derailment
  7. [7] Chang & Danescu-Niculescu-Mizil. Trouble on the Horizon: Forecasting the Derailment of Online Conversations as they Develop, EMNLP-IJCNLP 2019 (arXiv:1909.01362) (2019) [archived]
    jonathan p. chang , cristian danescu-niculescu-mizil
    ; Cheng, Bernstein, Danescu-Niculescu-Mizil & Leskovec. Anyone Can Become a Troll: Causes of Trolling Behavior in Online Discussions, CSCW 2017 (2017) [archived]
    justin cheng, michael bernstein, cristian danescu-niculescu-mizil, jure leskovec
  8. [8] Cheng, Bernstein, Danescu-Niculescu-Mizil & Leskovec. Anyone Can Become a Troll: Causes of Trolling Behavior in Online Discussions, CSCW 2017 (2017) [archived]
    a predictive model of trolling behavior shows that mood and discussion context together can explain trolling behavior better than an individual's history of trolling