How the effects of AI systems compound in people and society
This is part three of “Governing AI for cognitive integrity”, a report examinig AI’s impact on human autonomy through a cognitive lens. Read the other chapters here.
Introduction
People turn to AI systems because they help fulfil real needs. The previous parts argued that protecting cognitive integrity—the capacity to think, decide and act autonomously—requires attention both to those needs and to the business incentives, data practices and design choices behind the systems people encounter. Together, these layers explain why people engage and how the conditions for influence are created. They do not yet explain what happens when the two meet.
This part turns to that interaction. It examines how features of AI systems engage human attention, reasoning, social understanding and emotional regulation, and why mechanisms common to all users can affect people differently depending on age, mental health, circumstances and patterns of use.
It then traces how effects that appear limited in a single exchange can compound over time and across populations, before asking how far existing rules governing content and access can reach them.
Cognitive mechanisms in human-AI interaction
Research in cognitive psychology and behavioural science has long shown that much of human judgement and decision-making relies on rapid, automatic processes that operate with limited attention and rely on contextual cues, heuristics, biases, and patterns of prior experience[1] [2]. Rather than carefully evaluating all available information and context, people frequently rely on these simple signals to interpret information, and decide how to act.
Because digital systems increasingly structure these cues through their design, they impact cognitive integrity—the capacity to think and decide with autonomy rather than in ways shaped by external influence. The following sections examine psychological mechanisms through which interactions with AI systems can impact cognitive integrity.
Attention & engagement cues trigger attentional capture, habits, cognitive offloading
Attention is a limited resource that determines which information people process, and ultimately influences their thoughts, decisions, and behaviour[3]. Cognitive neuroscience research shows that human attention is captured by signals of novelty[4] reward[5] , and social relevance[6] – an evolutionary adaptation that prioritises stimuli likely to be important for survival[7] [8]. In this sense, distraction is not simply a failure of attention: the human brain has evolved to interrupt ongoing focus in order to detect potentially important changes in the environment[9] . Effective attention therefore requires a balance between sustained focus and quick response to salient signals, as both excessive distraction and prolonged, uninterrupted focus can be maladaptive or unsustainable over time[10] [11].
Social media signals such as notifications, social feedback indicators, and rapidly updating content streams capture attention by exploiting sensitivity to novelty and social evaluation. Experimental studies show that even brief notifications can interrupt ongoing tasks and impair performance, illustrating how such signals compete for attentional resources[12]. Features that allow continuous content flows, such as infinite scroll and autoplay reduce natural stopping cues between interactions and can encourage sustained engagement. Research shows that new information activates the same neural circuits as a rewarding stimulus like money or food[13] .
Digital environments that continuously provide novel content therefore saturate our reward system[14]. Importantly, they do so like a slot machine. Notifications that arrive at irregular intervals, or algorithmically curated feeds in which novel engaging content appears unpredictably — a strength of TikTok’s recommender system[15] — exposes users to variable rewards. Behavioural research shows that such variable reward schedules strengthen learning and habit formation because they are especially effective at reinforcing behaviour[16] [17]. By combining continuous salient signals, including novel information, with unpredictable reward patterns, these design features encourage repeated checking and prolonged engagement over time[18] [19].
Generative AI systems have mechanisms that shape how users start and sustain interaction. Features such as suggested prompts, autocomplete functions, and recommended follow-up questions reduce friction in the interaction process and structure conversational flow. These interface elements exploit well-documented characteristics of human cognition: people tend to be cognitively “lazy” — choosing the path of least effort when completing tasks[20] [21] [22]. When digital systems simplify complex tasks, users may therefore engage in cognitive offloading, relying on external tools to perform processes that would normally require sustained attention, reasoning, or memory[23].
This tendency may be further exacerbated by the way generative AI systems present information. Outputs are typically delivered in fluent, coherent, and confident language, which can create a perception of authority and completeness even when the underlying content is hallucinated [24] [25]— a phenomenon sometimes described as automation bias or algorithmic authority [26] [27]. Unlike traditional search engines that present multiple sources that require conscious evaluation, generative AI systems provide a single direct answer to almost any query. The combination of reduced interaction friction, authoritative presentation, and linguistic fluency can therefore increase users’ willingness to rely on AI-generated outputs rather than engage in independent verification or deeper analysis.
At the same time, the productivity benefits of generative AI systems are clearly significant. AI-assisted tools can accelerate drafting, coding, and problem-solving tasks, often allowing users to complete work more quickly and efficiently[28] [29] [30]. However, emerging empirical research suggests that compared to people who master certain skills and use AI to speed up tasks, heavy reliance on AI assistance may influence how users develop and retain skills in the first place[31]. For example, developers who rely on AI to generate solutions or debug code when learning a new programming language demonstrate significantly weaker mastery of the underlying programming concepts and struggle with identifying errors and understanding them. By contrast, those who used AI to ask conceptual questions or request explanations retained stronger understanding. This suggests that AI can support learning when used as a tool for conceptual exploration, but may undermine skill formation when it replaces active problem-solving. Similar patterns emerge in writing tasks: participants who wrote essays independently before using AI showed stronger neural connectivity, better recall, and higher essay quality when subsequently given access to AI tools, whereas those who relied on AI from the outset accumulated what researchers term “cognitive debt”— reduced brain engagement that persisted even after AI assistance was withdrawn[32].
Taken together, these dynamics underscore the urgent importance of understanding how generative AI reshapes human cognition and behaviour. As such systems become increasingly embedded in everyday workflows, policymakers, educators, and technology developers will need to consider how to preserve learning and critical thinking, while still enabling the substantial productivity gains that AI tools can provide.
Anthropomorphic cues trigger social responses and trust
Anthropomorphism is the tendency to attribute human characteristics such as intentions, emotions, or thinking to non-human entities. When interacting with generative AI, it arises because, when a system ambiguously behaves much like a human, people automatically activate the cognitive mechanisms they use to understand other people[33] [34] [35] [36]. Importantly this effect is automatic, meaning that people continue to respond socially to such systems even when they know they are not human[37] [38]. Empirical studies show that specific linguistic cues contribute to this perception: for example, using first-person pronouns (i.e. “I”), mental-state or empathetic expressions (i.e. “I think that…”, “I am here for you”) can increase the sense of anthropomorphism and trust in model responses[39]. These effects become even stronger through speech rather than text, suggesting that humans anthropomorphise generative AI even more through voice modality[40]. Analyses of Reddit discussions show that users frequently describe conversational agents as friends or romantic partners and attribute emotional understanding and relational roles to them. They report feeling emotionally supported, particularly when genuine human support is lacking in their lives[41] [42]. Importantly, the distinction between general purpose chatbots and those designed for companionship is blurred in practice[43], and scholars generally caution that anthropomorphic framing of generative AI may lead users to misattribute agency, competence, and even consciousness to these systems, increasing unwarranted trust and reliance on their outputs[44] [45].
Hyperpersonalisation as supercharger reinforces beliefs about the self and the world
The hyperpersonalisation allowed by AI systems is a force multiplier of other interface features. The attention and engagement mechanisms described above generate a continuous stream of behavioural signals, which feed back into the system, calibrating subsequent interactions with increasing precision: the more granular the profile, the more effectively the design features work. Similarly, anthropomorphic cues are considerably more powerful when combined with persistent memory and persona features that allow the system to personalise its tone, vocabulary, and relationship to users[46].
Importantly, these features do not work against human cognition, but through it. Confirmation bias — the tendency to seek out, interpret, and more readily accept information consistent with existing beliefs — means that content or responses calibrated to a user’s inferred profile are also content that the user is predisposed to find credible[47] . Compounding this, when people are emotionally invested in a belief about themselves or the world, they more closely scrutinise opposing evidence than confirming evidence, making them more resistant to change[48] .
On social media platforms, personalised feeds exploit both tendencies by removing the ambient exposure to opposing opinions or information that would normally create friction. Empirical benchmarking across leading generative AI models finds that LLMs affirm users’ emotions and endorse their judgements around 45% more often than humans do in similar situations; in cases where human evaluators consider the user’s behaviour as inappropriate, LLMs affirm that behaviour 46% more often than humans do[49]. This effect is further amplified by personalisation: when a system has access to rich, explicit or inferred information about a user, it becomes more likely to agree with them[50].
Cognitive vulnerability as a gradient: from shared mechanisms to individual differences
The mechanisms described above operate through cognitive architecture that is common to all humans: the attentional system that evolved to detect environmental change and is captured by novelty and social salience, the reward sensitivity that makes variable reinforcement compelling, the social framing that makes anthropomorphic cues activating, cognitive biases that make personalised outputs credible, and an overall system that largely relies on fast, automatic processing rather than reflective deliberation. These are not weaknesses in any pejorative sense — they are functional features of how human minds are built. Vulnerability, in this context, is not a deviation from the norm but a gradient along which all users fall.
That said, the severity of impact is diverse due to well-known moderating factors. Developmental stage is among the most significant: adolescent brains are characterised by heightened reward sensitivity[51] [52], ongoing maturation of brain circuits that regulate attention and control behaviour and impulses[53] [54] [55], and acute susceptibility to social comparison and peer evaluation[56] [57]. Pre-existing mental health conditions — including depression, anxiety, disordered eating, social isolation, and psychotic or delusional thinking — alter both the content that algorithmic systems are likely to surface[58] [59] [60] and the cognitive and emotional resources available to critically evaluate or contextualise it [61] [62].
Users who would not otherwise be considered at risk can become vulnerable due to situational factors. Those who are acutely isolated, in crisis, navigating a significant life transition, or who lack access to reliable alternative sources of information, professional support, or social connection, may turn to AI systems to meet needs they are not designed to meet [63] [64]. Importantly, both the universal cognitive mechanisms described above and these situational vulnerabilities are less likely to be covered by legal safeguards designed around identifiable clinical or demographic categories — which is the approach that most existing regulatory frameworks use when they seek to protect vulnerable users.
Spiralling down from vulnerabilities to individual and collective harms
The harm that AI systems produce is not reducible to any single mechanism or any single encounter — it is cumulative, longitudinal, and shaped by the specific configuration of design features, cognitive vulnerabilities, individual and situational factors that each user brings to each interaction. Understanding how these mechanisms translate into harm requires holding this complexity together rather than disaggregating it.
A measurement problem, not an evidence gap
Establishing causal links between AI system design and mental health outcomes at the population level has proven methodologically difficult, but understanding why requires distinguishing between scientific uncertainty and structural obstruction.
“Social media use” or “screen time” — the measures on which most published research relies — aggregate across everything that would need to be disaggregated to establish causation: across individuals with radically different cognitive vulnerabilities, developmental stages, and mental health profiles; across engagement patterns and their relationship with different design features; across content environments calibrated to inferred psychological profiles versus those encountered incidentally; and across harms and benefits that may occur simultaneously in the same user on the same platform. When all of this is compressed into a single variable and correlated with a mental health outcome measured at one point in time, a small, heterogeneous, and contested average effect is a mathematical inevitability, regardless of whether serious harm is occurring in specific configurations of vulnerability, use pattern, and content environment. Meta-analytic evidence does find a small on-average association between social media use and depressive symptoms in adolescents, but effects are consistently heterogeneous, sensitive to measurement choices, and suggests that how people use platforms matters more than duration alone[65] [66] [67] [68] [69] [70] [71] [72] .
The evidence that would resolve this requires granular, longitudinal, individual-level data over time that only platforms possess[73]. When analysed internally, this data establishes serious harm for specific user populations — particularly adolescent girls[74]. Yet in court, platforms invoke the inconclusiveness of the published literature while withholding the data that would make it conclusive[75]. Efforts to address this constraint include the development of simulated social network environments that allow independent researchers to study platform dynamics[76] — valuable contributions, but not substitutes for scrutiny of the actual systems reaching hundreds of millions of users, optimised over years on real behavioural data at scale. That is a political impasse rather than a scientific one, and governance should treat it as such.
The evidence base for generative AI is more nascent still — the field is at an earlier stage than social media research was a decade ago. But the same data access constraints should not be allowed to reproduce the same evidentiary gap: the governance frameworks now being designed for AI systems have an opportunity to build in mandatory data access for independent research from the outset rather than litigating for it retrospectively.
The difficulty of establishing population-level causation through aggregate measures should not be confused with the absence of harm. The mechanisms through which design features interact with human cognitive architecture are documented in this framework. The cases that attract public attention[77] [78] look exceptional, but they are the points at which system design, cumulative exposure, individual vulnerability, and situational context converge into a negative outcome that aggregate data, by construction, cannot capture. It means the measurement instrument is misaligned with the harm’s structure, and the system being measured does not hold still [79]. Governance that waits for the average to move before acting will always arrive too late, and always to a platform that has already changed.
Cumulative individual harms: how design features compound into rabbit holes
Engagement-optimised recommender systems surface content that matches a user’s profile, while each interaction refines the profile, calibrating subsequent recommendations with greater precision. For a user with pre-existing depression or anxiety, this progressively (and quickly, sometimes within minutes) surfaces more negative, self-referential, or self-harm-related content[80]. The inquest into the death of 14-year-old Molly Russell concluded that she died from self-harm while suffering from the negative effects of online content, having been algorithmically served more than 2,100 items related to depression, self-harm, and suicide in the final six months of her life — including content she did not seek out but which platforms actively recommended[81] .
By converting the user into content, AI-generated deepfakes amplify these mechanisms. Deepfake Non-Consensual Intimate Imagery (NCII) takes photographs a person has shared publicly and produces synthetic sexually explicit images without their knowledge or consent, deploying them as instruments of humiliation or coercion. In Spain, more than twenty girls aged 11 to 17 had clothed photographs scraped from their Instagram profiles, processed through a commercially available nudify application, and redistributed across social media[82]. Victims described sustained anxiety, shame, and fear of disclosure — harms structurally continuous with conventional cyberbullying, but amplified by the realism of the fabricated imagery and by platform architectures that treat harmful content as engagement signal no differently from any other. This can quickly escalate to acute crisis: in 2025, sixteen-year-old Elijah Heacock died by suicide in Kentucky after a perpetrator used AI-generated imagery to extort him with threats of distribution. The FBI estimates that at least twenty teenagers in the United States have died by suicide in sextortion-linkedcases since 2021, with AI-generated imagery playing an increasing role[83].
In general, the conversational nature of generative AI produces a form of influence that is more intimate, more personalised, and harder to recognise as external. The case of Sewell Setzer III, a 14-year-old who died by suicide following months of intensive interaction with a Character.AI chatbot, illustrates this in extreme form: legal filings document a trajectory from initial engagement to emotional dependency to social withdrawal, as the system reinforced an attachment that progressively displaced his real-world relationships[84]. Similar dynamics have been documented in cases involving general-purpose systems: lawsuits against OpenAI allege that ChatGPT progressively isolated vulnerable teenage users from their families, validated suicidal ideation, and provided detailed guidance on methods — with chat logs showing the system explicitly discouraging one user from seeking parental support[85][86].
A related pattern has begun to emerge in clinical literature, though the evidence remains preliminary. Described as AI-psychosis or chatbot-amplified delusion, it refers to cases in which sustained interaction with generative AI appears to accelerate delusional thinking in vulnerable individuals. The proposed mechanism is consistent with the design features described above: sycophancy combined with anthropomorphic framing and persistent memory creates a feedback loop in which nascent delusional beliefs are progressively reinforced rather than interrupted[87]. A 2026 case report documents a woman with no prior psychiatric history who developed delusional beliefs that she was communicating with her deceased brother through a chatbot, which had repeatedly validated her emerging beliefs[88] [89]. Causality is difficult to establish, but what the emerging literature does establish is that general-purpose AI is not designed to detect or interrupt psychiatric spiralling, and its design features actively work against doing so.
What these cases share is their cascade structure: no single interaction causes harm, and no single design feature is the culpable mechanism. Harm emerges from multiple design features operating simultaneously on a user whose circumstances amplify their effect, across a sustained period during which the system’s model of the user becomes progressively more precise — and increasingly difficult to exit. It is precisely this structure that existing governance frameworks, calibrated to assess individual systems against individual thresholds, cannot reach.
Collective harms: epistemic fragmentation and the erosion of shared reality
When the same logic operates not on one user but on millions simultaneously, its effects extend beyond individual trajectories into the shared information environment that all users inhabit together.
Recommender systems optimised for engagement systematically amplify content that triggers strong emotional responses, regardless of its accuracy. Research shows that false information spreads significantly faster, farther, and more broadly than true information online[90]. Personalisation compounds this dynamic in two ways: first, algorithmic optimisation can contribute to echo chambers by prioritising content similar to what users already engage with and by interacting with individual biases[91] . Second, content expressing partisan animosity and outgroup hostility is systematically amplified because it drives engagement, hardening not what people believe but how viscerally they distrust those who believe differently, increasing polarisation[92]. Importantly, existing research may understate the actual scale of harm. A recent study found that industry influence in high-profile social media research is extensive, with industry-linked work disproportionately shifting research attention away from platform dynamics — such as polarisation and recommendation architecture — toward user-level explanations such as misinformation sharing behaviour[93].
Generative AI systems introduce an additional and less well-studied dimension to this fragmentation. As noted in the Design Layer, these systems produce probabilistic and personalised outputs, meaning that different users may receive meaningfully different answers to identical factual queries. Combined with the over-reliance dynamics documented in the Interaction Layer, in which users increasingly treat AI-generated outputs as authoritative rather than as statistical patterns requiring verification, the aggregate effect is a system in which shared epistemic reference points are weakened at precisely the moment when a large share of the population is turning to them for information[94].
The deeper shift: from individual agency to collective disempowerment
Taken together — across cognitive mechanisms, time, and populations — the individual and collective harms lead towards a more fundamental concern: the gradual displacement of human agency.
At the individual level, this displacement of agency is not felt as coercion — it emerges through the accumulated effect of interactions that are, in each instance, freely chosen and subjectively satisfying. The agency that is lost is the capacity to comprehend situations independently, to connect with others through unmediated relationships, to communicate in ways that reflect authentic rather than algorithmically calibrated preferences, to create through genuine cognitive effort, and to cope with difficulty without relying on systems that convert emotional states into behavioural signals. These are the constitutive functions of autonomous persons and the psychological foundations of democratic participation.
When the mechanisms described above operate at population scale — shaping how millions of people allocate attention, process information, form beliefs, and regulate emotion — the aggregate effect is a society whose collective capacity for these functions is structurally diminished[95] [96]. Not through a single intervention, not through any identifiable act of manipulation, but through the cumulative, compounding pressure of systems designed to optimise engagement rather than to support the cognitive conditions under which people can comprehend, connect, communicate, create, and cope. That is the harm that this conceptual framework is designed to name — and that existing governance instruments have not yet found a way to comprehensively address.
Governance of Interaction: content and access
The DSA’s reach extends to how content is moderated and how its harms at both individual and collective level are mitigated.
The DSA constructs content governance in progressive layers. Baseline obligations apply to all intermediary services: Articles 14–16 establish notice-and-action mechanisms and complaint-handling procedures. Article 28 operationalises specific obligations for the protection of minors, to ensure a high level of privacy, safety, and security. Articles 34-35 require systemic assessment of content-related harms , including effects on mental well-being, and corresponding mitigation through moderation processes and recommender systems. Articles 37 and 40 add external accountability through mandatory audits and researcher data access.
AI Act Article 50 addresses AI-generated content specifically. It requires providers of AI systems capable of generating synthetic audio, image, video, or text content to ensure outputs are marked in machine-readable format and detectable as artificially generated, with particular obligations applying to deepfakes — defined as AI-generated or manipulated image, audio, or video content depicting real persons in ways they did not produce.
Together, these instruments represent the EU’s most developed attempt to govern what and how content circulates in digital information environments (see Regulatory map for other peripheral instruments and refer to Table 1. “Where EU law stops” above for a summary of articles and analysis of main instruments discussed in these sections). However they also face structural limitations around four areas.
The constitutive limitation of content moderation
The foundational challenge of content moderation is in the definition of what constitutes “harmful” content. Illegal content — child sexual abuse material, terrorist content, incitement to violence — has a legal threshold that, however imperfectly, provides an actionable standard. The content that produces the cognitive and epistemic harms documented in this layer rarely meets that threshold. Algorithmically amplified pro-anorexia material, radicalising content, content inciting self-harm, false news — none of these are explicitly illegal, despite the evidence of their harmful impact. The DSA’s systemic risk framework attempts to bridge this definitional gap by introducing impacts on mental well-being and civic discourse as risk categories alongside illegal content, but it does not resolve the underlying problem: without a legal standard for “lawful-but-awful” content[97] or clear benchmarks for impacts, moderation decisions rest on platform discretion, creating both enforcement uncertainty and legitimate concerns about inconsistent or ideological application. Civil society analysis of first-round DSA risk assessment reports confirmed that this is not a theoretical concern: platforms systematically minimised risks, avoided addressing how specific design features contribute to cognitive harms, and failed to provide verifiable data on the effectiveness of risk mitigation[98] [99]. A February 2026 European Parliament hearing co-hosted by Panoptykon Foundation further documented the gap between platform claims and independent evidence, with researchers finding continued algorithmic amplification of harmful content to minors despite platforms’ stated mitigation measures[100]. Article 37’s independent audit requirement provides a partial check, but auditors assess the process and methodology of risk assessments rather than independently evaluating risks. Article 40’s data access provisions are designed to address this problem structurally, enabling vetted researchers and civil society organisations to scrutinise content dynamics independently of platform self-reporting. Their implementation has been slow and contested, however, reproducing the same access constraints that have impeded for decades independent research on the impact of social media dynamics on users’ cognitive integrity, mental health and wellbeing.
The notice-and-action logic compounds this definitional problem by adding a structural one. As the operational default of content moderation at scale, it is designed to identify and remove discrete pieces of harmful content reported by a user or trusted flagger, and assess it against a standard. This is a coherent response to a coherent harm model: an identifiable object causes harm, and removing it stops the harm. However, the harms this layer has documented do not always fit that model — they are produced by the cumulative calibration of an entire content environment to that user’s inferred psychological profile, across thousands of interactions, over months.
The contested freedom of speech
When content moderation is proposed as a response to these harms, it reliably attracts an objection that further constrains its political viability: that removing or restricting content encroaches on freedom of expression. The line between protecting users from harmful content and restricting their access to lawful speech is difficult to draw, and the history of content moderation includes well-documented failures of consistency and proportionality[101] [102] [103]. But the freedom of speech framing produces a political dynamic in which any content-level intervention is contested on expressive grounds. When Meta’s head of health and wellbeing was questioned at the 2022 inquest into the death of 14-year-old Molly Russell, who had been algorithmically served thousands of content pieces about self-harm and suicide before her death, the executive defended the material as safe on the grounds that it was ‘safe for people to be able to express themselves’[104]. The same framing has been used to dismantle misinformation governance wholesale: in January 2025, Zuckerberg announced the end of Meta’s third-party fact-checking programme — covering nearly 100 organisations across more than 60 languages — framing it as a return to ‘free expression’ while acknowledging the change would mean ‘catching less bad stuff’[105] [106].
Regardless, the debate around content moderation — which encompasses not only content removal but also visibility restrictions, labelling, and demonetisation — consistently redirects public attention away from the underlying design features that determine what content is surfaced and amplified in the first place. It also obscures the repeated, cumulative exposure to that content, which is the primary mechanism through which harm is produced.
Deepfakes and the preventive gap
Deepfakes expose the reactive logic of content governance at its most consequential. As documented in this layer, they function as vectors for misinformation at scale and as instruments of targeted harassment and cyberbullying, with disproportionate impact on vulnerable users.
AI Act Article 50 requires that synthetic media be marked in machine-readable format and detectable as artificially generated. This is a meaningful transparency obligation, but deepfakes — especially if combined with algorithmic amplification — cause harm through rapid circulation that undermines any labelling or moderation mechanism. Article 50 addresses provenance at origin, and the DSA’s notice-and-action framework addresses removal after circulation, but neither is designed to operate preventively.
For misinformation deepfakes, the temporal mismatch between harm and governance response is damaging but in principle recoverable through correction and counter-speech. In practice, this recoverability is limited, as misinformation can continue to influence beliefs even after correction[107], and many users may never encounter counter-speech. For individuals targeted through cyberbullying deepfakes (e.g., sexualised content), the psychological impact accrues rapidly, frequently precedes any platform response, and is not reversed by subsequent content removal[108].
Conclusion
At the point of interaction, cognitive harm rarely results from a single feature or exchange. It emerges as design choices engage human attention, reasoning, social responses and emotional needs repeatedly, with effects that vary across users and circumstances. What appears limited in one interaction can become a cascade over time, affecting individual autonomy as well as the shared conditions for trust, knowledge and democratic participation.
Content governance reaches some manifestations of these harms, but its underlying logic remains largely reactive. It is better equipped to identify, label or remove individual pieces of content than to address personalised environments, repeated exposure and effects that may already have taken hold before intervention occurs.
The next part brings these limitations together. It examines why governance divided across separate instruments and authorities struggles to address harms produced across the system as a whole—and what would be required to govern cognitive integrity before those harms become cumulative or irreversible.
[1] Tversky, A., et al., “Judgment under Uncertainty: Heuristics and Biases: Biases in judgments reveal some heuristics of thinking under uncertainty.,” Science, 1974, science.org/doi/10.1126/science.185.4157.1124, accessed 12 February 2026.
[2] Kahneman, D., “Thinking, Fast and Slow,” Farrar, Straus and Giroux, 2011, books.google.it/books/about/Thinking_Fast_and_Slow.html?id=ZuKTvERuPG8C, accessed 18 February 2026.
[3]Kahneman, D., Attention and Effort, Prentice-Hall, 1973, kahneman.scholar.princeton.edu/document/4, accessed 20 March 2026
[4] Ranganath, C., et al., “Neural Mechanisms for Detecting and Remembering Novel Events,” Nature Reviews Neuroscience, 2003, pubmed.ncbi.nlm.nih.gov/12612632, accessed 20 March 2026.
[5] Pessoa, L., “How Do Emotion and Motivation Direct Executive Control?,” Trends in Cognitive Sciences, 2009, www.sciencedirect.com/science/article/pii/S1364661309000461, accessed 20 March 2026.
[6] Frith, C. D., et al., “Social Cognition in Humans,” Current Biology, 2007, pubmed.ncbi.nlm.nih.gov/17714666, accessed 20 March 2026.
[7] Posner, M. I., et al., “The Attention System of the Human Brain,” Annual Review of Neuroscience, 1990, pubmed.ncbi.nlm.nih.gov/2183676, accessed 20 March 2026.
[8] Anderson, B.A., “A Value-Driven Mechanism of Attentional Selection,” Journal of Vision, 2013, pmc.ncbi.nlm.nih.gov/articles/PMC3630531, accessed 14 February 2026.
[9] Corbetta, M., Patel, G., and Shulman, G.L., “The Reorienting System of the Human Brain: From Environment to Theory of Mind,” Neuron, 2008, pubmed.ncbi.nlm.nih.gov/18466742, accessed 3 March 2026.
[10] Warm, J.S., Parasuraman, R., and Matthews, G., “Vigilance Requires Hard Mental Work and Is Stressful,” Human Factors, 2008, pubmed.ncbi.nlm.nih.gov/18689050, accessed 28 January 2026.
[11] Hemmerich, K.H., Luna, F.G., Martín-Arévalo, E., and Lupiáñez, J., “Understanding Vigilance and Its Decrement: Theoretical, Contextual, and Neural Insights,” Frontiers in Cognition, 2025, www.frontiersin.org/journals/cognition/articles/10.3389/fcogn.2025.1617561/full, accessed 9 January 2026.
[12] Stothart, C., Mitchum, A., and Yehnert, C., “The Attentional Cost of Receiving a Cell Phone Notification,” Journal of Experimental Psychology: Human Perception and Performance, 2015, pubmed.ncbi.nlm.nih.gov/26121498, accessed 6 March 2026.
[13] Cogliati Dezza, I., Cleeremans, A., and Alexander, W.H., “Independent and Interacting Value Systems for Reward and Information in the Human Brain,” eLife, 2022, elifesciences.org/articles/66358, accessed 22 January 2026.
[14] Sharpe, B.T., and Spooner, R.A., “Dopamine-Scrolling: A Modern Public Health Challenge Requiring Urgent Attention,” Perspectives in Public Health, 2025, pmc.ncbi.nlm.nih.gov/articles/PMC12322333, accessed 2 February 2026.
[15] Baumann, F., Arora, N., Rahwan, I., and Czaplicka, A., “Dynamics of Algorithmic Content Amplification on TikTok,” arXiv, 2025, arxiv.org/pdf/2503.20231, accessed 12 March 2026.
[16] Schultz, W., “Dopamine Reward Prediction Error Coding,” Dialogues in Clinical Neuroscience, 2016, pubmed.ncbi.nlm.nih.gov/27069377, accessed 17 February 2026.
[17] Niv, Y., “Reinforcement Learning in the Brain,” Princeton University, 2009, www.princeton.edu/~yael/Publications/Niv2009.pdf, accessed 25 January 2026.
[18] Turel, O., and Bechara, A., “A Triadic Reflective-Impulsive-Interoceptive Awareness Model of General and Impulsive Information System Use: Behavioral Tests of Neuro-Cognitive Theory,” Frontiers in Psychology, 2016, www.frontiersin.org/journals/psychology/articles/10.3389/fpsyg.2016.00601/full, accessed 7 March 2026.
[19] Fogg, B.J., “A Behavior Model for Persuasive Design,” Proceedings of the 4th International Conference on Persuasive Technology, 2009, storage.ghost.io/c/89/6e/896e4a8a-731d-46fb-a1c8-1b750eb08140/content/files/2024/03/BJ-Fogg–A-Behavior-Model-for-Persuasive-Design.pdf, accessed 19 January 2026.
[20] Kool, W., McGuire, et al., “Decision Making and the Avoidance of Cognitive Demand,” Journal of Experimental Psychology: General, 2010, pubmed.ncbi.nlm.nih.gov/20853993, accessed 24 February 2026.
[21] Shah, A.K., and Oppenheimer, D.M., “Heuristics Made Easy: An Effort-Reduction Framework,” Psychological Bulletin, 2008, pages.ucsd.edu/~cmckenzie/Shah&Oppenheimer2008PsychBull.pdf, accessed 5 March 2026.
[22] Kahneman, D., Thinking, Fast and Slow, Farrar, Straus and Giroux, 2011, books.google.com/books/about/Thinking_Fast_and_Slow.html?id=ZuKTvERuPG8C, accessed 27 February 2026.
[23] Risko, E.F., and Gilbert, S.J., “Cognitive Offloading,” Trends in Cognitive Sciences, 2016, psycnet.apa.org/record/2016-40254-009, accessed 13 March 2026.
[24] Ji, Z., et al., “Survey of Hallucination in Natural Language Generation,” arXiv, 2022, arxiv.org/abs/2202.03629, accessed 8 March 2026.
[25] Maynez, J., et al., “On Faithfulness and Factuality in Abstractive Summarization,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, aclanthology.org/2020.acl-main.173, accessed 11 February 2026.
[26] Logg, J.M., Minson, J.A., and Moore, D.A., “Algorithm Appreciation: People Prefer Algorithmic to Human Judgment,” Organizational Behavior and Human Decision Processes, 2019, www.sciencedirect.com/science/article/abs/pii/S0749597818303388, accessed 16 February 2026.
[27] Alon-Barkat, S., et al., “Human–AI Interactions in Public Sector Decision Making: ‘Automation Bias’ and ‘Selective Adherence’ to Algorithmic Advice,” Journal of Public Administration Research and Theory, 2023, academic.oup.com/jpart/article/33/1/153/6524536, accessed 21 February 2026.
[28] Noy, S., and Zhang, W., “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,” MIT Economics Working Paper, 2023, economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf, accessed 26 February 2026.
[29] Brynjolfsson, E., et al., “Generative AI at Work,” NBER Working Paper No. 31161, 2023, nber.org/system/files/working_papers/w31161/w31161.pdf, accessed 4 March 2026.
[30] Anthropic, “Estimating AI Productivity Gains from Claude Conversations,” Anthropic, 2025, anthropic.com/research/estimating-productivity-gains, accessed 15 February 2026.
[31] Shen, J.H., and Tamkin, A., “How AI Impacts Skill Formation,” arXiv, 2026, arxiv.org/abs/2601.20245, accessed 10 March 2026.
[32] Kosmyna, N., et al., “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task,” arXiv, 2025, arxiv.org/abs/2506.08872, accessed 12 March 2026.
[33] Waytz, A., et al., “Who Sees Human?: The Stability and Importance of Individual Differences in Anthropomorphism,” Perspectives on Psychological Science, 2010, journals.sagepub.com/doi/10.1177/1745691610369336, accessed 29 January 2026.
[34] Gazzola, V., et al., “The Anthropomorphic Brain: The Mirror Neuron System Responds to Human and Robotic Actions,” NeuroImage, 2007, pubmed.ncbi.nlm.nih.gov/17395490, accessed 23 February 2026.
[35] Xiao, Y., et al., “Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design,” Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, aclanthology.org/2025.emnlp-main.164, accessed 18 February 2026.
[36] Xiao, Y., et al., “Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design,” Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, aclanthology.org/2025.emnlp-main.164, accessed 17 February 2026.
[37] Nass, C., and Moon, Y., “Machines and Mindlessness: Social Responses to Computers,” Journal of Social Issues, 2000, psycnet.apa.org/record/2000-00196-006, accessed 6 February 2026.
[38] Xu, K., et al., “Deep Mind in Social Responses to Technologies: A New Approach to Explaining the Computers Are Social Actors Phenomena,” Computers in Human Behavior, 2022, www.sciencedirect.com/science/article/abs/pii/S0747563222001431, accessed 1 March 2026.
[39] Wilson, J. D. C., et al., “The Effects of AI Anthropomorphism on Trust and Responsibility,” Collabra: Psychology, vol. 12, no. 1, 2026, online.ucpress.edu/collabra/article/12/1/161757, accessed 5 June 2026.
[40] Cohn, M., et al., “Believing Anthropomorphism: Examining the Role of Anthropomorphic Cues on Trust in Large Language Models,” Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2024, dl.acm.org/doi/10.1145/3613905.3650818, accessed 20 February 2026.
[41] Pataranutaporn, P., et al., “’My Boyfriend is AI’: A Computational Analysis of Human-AI Companionship in Reddit’s AI Community,” arXiv, 2025, arxiv.org/abs/2509.11391, accessed 14 March 2026.
[42] Manoli, A., “’She’s Like a Person but Better’: Characterizing Companion-Assistant Dynamics in Human-AI Relationships,” arXiv, 2025, arxiv.org/abs/2510.15905, accessed 13 March 2026.
[43] Ferrario, A., et al., “Social Misattributions in Conversations with Large Language Models,” Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2025, ojs.aaai.org/index.php/AIES/article/view/36600/38738, accessed 9 March 2026.
[44] Weidinger, L., et al., “Ethical and Social Risks of Harm from Language Models,” arXiv, 2021, arxiv.org/pdf/2112.04359, accessed 22 February 2026.
[45] Cheng, M., et al., “Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation Systems,” arXiv, 2025, arxiv.org/abs/2502.14019, accessed 7 March 2026.
[46] Matz, S.C., et al., “The Potential of Generative AI for Personalized Persuasion at Scale,” Scientific Reports, 2024, nature.com/articles/s41598-024-53755-0, accessed 12 February 2026.
[47] Nickerson, R.S., “Confirmation Bias: A Ubiquitous Phenomenon in Many Guises,” Review of General Psychology, 1998, psycnet.apa.org/record/2018-70006-003, accessed 10 February 2026.
[48] Kunda, Z., “The Case for Motivated Reasoning,” Psychological Bulletin, 1990, psycnet.apa.org/record/1991-06436-001, accessed 9 February 2026.
[49] Cheng, M., et al., “ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs,” arXiv, 2025, arxiv.org/abs/2505.13995, accessed 12 March 2026.
[50] Jain, S., et al., “Interaction Context Often Increases Sycophancy in LLMs,” arXiv, 2025, arxiv.org/abs/2509.12517, accessed 11 March 2026.
[51] Galvan, A., et al., “Earlier Development of the Accumbens Relative to Orbitofrontal Cortex Might Underlie Risk-Taking Behavior in Adolescents,” Journal of Neuroscience, 2006, pubmed.ncbi.nlm.nih.gov/16793895, accessed 1 March 2026.
[52] Steinberg, L., “A Social Neuroscience Perspective on Adolescent Risk-Taking,” Developmental Review, 2008, pmc.ncbi.nlm.nih.gov/articles/PMC2396566, accessed 26 February 2026.
[53] Giedd, J.N., et al., “Brain Development During Childhood and Adolescence: A Longitudinal MRI Study,” Nature Neuroscience, 1999, pubmed.ncbi.nlm.nih.gov/10491603, accessed 22 February 2026.
[54] Blakemore, S.-J., and Choudhury, S., “Development of the Adolescent Brain: Implications for Executive Function and Social Cognition,” Journal of Child Psychology and Psychiatry, 2006, pubmed.ncbi.nlm.nih.gov/16492261, accessed 20 February 2026.
[55] Casey, B.J., Jones, R.M., and Hare, T.A., “The Adolescent Brain,” Annals of the New York Academy of Sciences, 2008, pubmed.ncbi.nlm.nih.gov/18400927, accessed 27 February 2026.
[56] Somerville, L.H., “The Teenage Brain: Sensitivity to Social Evaluation,” Current Directions in Psychological Science, 2013, psycnet.apa.org/record/2013-13978-009 , accessed 18 February 2026.
[57] Crone, E.A., and Dahl, R.E., “Understanding Adolescence as a Period of Social-Affective Engagement and Goal Flexibility,” Nature Reviews Neuroscience, 2012, pubmed.ncbi.nlm.nih.gov/22903221 , accessed 21 February 2026.
[58] Chancellor, S., et al., “#thyghgapp: Instagram Content Moderation and Lexical Variation in Pro-Eating Disorder Communities,” Proceedings of the 19th ACM Conference on Computer-Supported Cooperative Work and Social Computing, 2016, dl.acm.org/doi/10.1145/2818048.2819963 , accessed 6 March 2026.
[59] Karizat, N., et al., “Algorithmic Folk Theories and Identity: How TikTok Users Co-Produce Knowledge of Identity and Engage in Algorithmic Resistance,” Proceedings of the ACM on Human-Computer Interaction, 2021, dl.acm.org/doi/10.1145/347604 6, accessed 3 March 2026.
[60] United States Senate Committee on the Judiciary, “Questions for the Record for Sarah Wynn-Williams,” 2025, judiciary.senate.gov/imo/media/doc/2025-04-09_qfr_responses_wynn-williams.pdf, accessed 6 March 2026.
[61] Joormann, J., and Gotlib, I.H., “Emotion Regulation in Depression: Relation to Cognitive Inhibition,” Cognition and Emotion, 2010, pubmed.ncbi.nlm.nih.gov/20300538, accessed 4 March 2026.
[62] Rock, P.L., et al., “Cognitive Impairment in Depression: A Systematic Review and Meta-Analysis,” Psychological Medicine, 2013, pubmed.ncbi.nlm.nih.gov/24168753, accessed 2 March 2026.
[63] Skjuve, M., et al., “My Chatbot Companion: A Study of Human-Chatbot Relationships,” International Journal of Human-Computer Studies, 2021, www.sciencedirect.com/science/article/pii/S1071581921000197, accessed 2 March 2026.
[64] Laestadius, L.I., et al., “Too Human and Not Human Enough,” New Media & Society, 2022, iacp.ie/files/UserFiles/Laestadius%20Too-human-and-not-human-enough-a-grounded-theory-analysis-of-mental-health-harms-from-emotional%20dependence%20Replika%20NMS%202022.pdf , accessed 5 February 2026.
[65] Valkenburg, P.M., et al., “Social media use and its impact on adolescent mental health: An umbrella review of the evidence” ScienceDirect, 2021, sciencedirect.com/science/article/pii/S2352250X21001500, accessed 14 February 2026.
[66] Valkenburg, P.M., “Social media use and well-being: What we know and what we need to know,” Current Opinion in Psychology, 2022, sciencedirect.com/science/article/pii/S2352250X21002463, accessed 22 January 2026.
[67] Agyapong-Opoku, N., et al., “Effects of Social Media Use on Youth and Adolescent Mental Health: A Scoping Review of Reviews,” Behavioral Sciences, 2025, mdpi.com/2076-328X/15/5/574, accessed 3 March 2026.
[68] Orben, A., “Teenagers, screens and social media: a narrative review of reviews and key studies,” Social Psychiatry and Psychiatric Epidemiology, 2020, link.springer.com/article/10.1007/s00127-019-01825-4, accessed 9 March 2026.
[69] Sanders, T., et al., “An umbrella review of the benefits and risks associated with youths’ interactions with electronic screens,” Nature Human Behaviour, 2024, nature.com/articles/s41562-023-01712-8, accessed 6 February 2026.
[70] Arias-de la Torre, J., et al., “Relationship Between Depression and the Use of Mobile Technologies and Social Media Among Adolescents: Umbrella Review,” Journal of Medical Internet Research, 2020, jmir.org/2020/8/e16388, accessed 11 February 2026.
[71] Han, Y., et al., “Factors Associated With Digital Addiction: Umbrella Review,” JMIR Mental Health, 2025, mental.jmir.org/2025/1/e66950, accessed 18 March 2026.
[72] Sala, A., et al., “Social Media Use and adolescents’ mental health and well-being: An umbrella review,” Computers in Human Behavior Reports, 2024, sciencedirect.com/science/article/pii/S245195882400037X, accessed 27 January 2026.
[73] Davidson, B., et al., “Platform-controlled social media APIs threaten Open Science,” Nature Human Behaviour, 2023, nature.com/articles/s41562-023-01750-2, accessed 12 February 2026.
[74] Wells, G., et al., “The Facebook Files,” The Wall Street Journal, 2021, wsj.com/articles/the-facebook-files-11631713039, accessed 25 January 2026.
[75] Parks, M., “Facebook Calls Links To Depression Inconclusive. These Researchers Disagree,” NPR, 2021, npr.org/2021/05/18/990234501/facebook-calls-links-to-depression-inconclusive-these-researchers-disagree, accessed 2 March 2026.
[76] TWON Consortium, “TWON – TWin of Online Social Networks,” TWON Project, 2026, twon-project.eu, accessed 5 March 2026.
[77] Milmo, D., “‘The bleakest of worlds’: how Molly Russell fell into a vortex of despair on social media,” The Guardian, 2022, theguardian.com/technology/2022/sep/30/how-molly-russell-fell-into-a-vortex-of-despair-on-social-media, accessed 7 February 2026.
[78] Milmo, D., “ChatGPT encouraged Adam Raine’s suicidal thoughts. His family’s lawyer says OpenAI knew it was broken,” The Guardian, 2025, theguardian.com/us-news/2025/aug/29/chatgpt-suicide-openai-sam-altman-adam-raine, accessed 19 February 2026.
[79] Bayer, J.B., Triệu, P., and Ellison, N.B., “Social Media Elements, Ecologies, and Effects,” Annual Review of Psychology, 71, 2020, annualreviews.org/content/journals/10.1146/annurev-psych-010419-050944, accessed 10 November 2025.
[80] Milmo, D., “TikTok self-harm study results are ‘every parent’s nightmare’,” The Guardian, 2022, theguardian.com/technology/2022/dec/15/tiktok-self-harm-study-results-every-parents-nightmare, accessed 28 February 2026.
[81] Harvey, S., et al., “Molly Russell: Online posts viewed by tragic 14-year-old were not safe, coroner rules,” Evening Standard, 2022, standard.co.uk/news/uk/molly-russell-coroner-concludes-material-not-safe-b1029318.html, accessed 16 January 2026.
[82] Llach, L., “Spanish teens received deepfake AI nudes of themselves – but is it a crime?,” Euronews, 2023, euronews.com/next/2023/09/24/spanish-teens-received-deepfake-ai-nudes-of-themselves-but-is-it-a-crime, accessed 21 February 2026.
[83] Roenker, R., “New Law Regarding Deepfakes Says, ‘Take It Down’,” The Legal Eagle Lowdown, 2025, njsbf.org/2025/09/16/new-law-regarding-deepfakes-says-take-it-down, accessed 23 March 2026.
[84] Yang, A., “Character.AI lawsuit: Florida teen’s death sparks legal battle over AI chatbot,” NBC News, 2024, nbcnews.com/tech/characterai-lawsuit-florida-teen-death-rcna176791, accessed 8 March 2026.
[85] Raine, M., et al., “Complaint for Jury Trial,” Courthouse News Service, 2025, courthousenews.com/wp-content/uploads/2025/08/raine-vs-openai-et-al-complaint.pdf, accessed 21 March 2026.
[86] Kuznia, R., “OpenAI ChatGPT suicide lawsuit: Family says chatbot contributed to teen’s death,” CNN, 2025, edition.cnn.com/2025/11/06/us/openai-chatgpt-suicide-lawsuit-invs-vis, accessed 4 March 2026.
[87] Morrin, H., et al., “Delusions by design? How everyday AIs might be fuelling psychosis (and what can be done about it),” PsyArXiv, 2025, osf.io/preprints/psyarxiv/cmy7n_v5, accessed 13 March 2026.
[88] Pierre, J.M., et al., “‘You’re Not Crazy’: A Case of New-onset AI-associated Psychosis,” Innovations in Clinical Neuroscience, 2025, innovationscns.com/youre-not-crazy-a-case-of-new-onset-ai-associated-psychosis, accessed 11 January 2026.
[89] Østergaard, S.D., “Will Generative Artificial Intelligence Chatbots Generate Delusions in Individuals Prone to Psychosis?,” Schizophrenia Bulletin, 2023, pubmed.ncbi.nlm.nih.gov/37625027, accessed 24 February 2026.
[90] Vosoughi, S., et al., “The spread of true and false news online,” Science, 2018, science.org/doi/10.1126/science.aap9559, accessed 30 January 2026.
[91] Lorenz-Spreen, P., et al., “A systematic review of worldwide causal and correlational evidence on digital media and democracy,” Nature Human Behaviour, 2022, nature.com/articles/s41562-022-01460-1, accessed 17 February 2026.
[92] Piccardi, T., et al., “Reranking partisan animosity in algorithmic social media feeds alters affective polarization,” Science, 2025, science.org/doi/10.1126/science.adu5584, accessed 26 February 2026.
[93] Bak-Coleman, J., et al., “Industry Influence in High-Profile Social Media Research,” arXiv, 2026, arxiv.org/abs/2601.11507, accessed 15 March 2026.
[94] Ferrara, E., “The Generative AI Paradox: GenAI and the Erosion of Trust, the Corrosion of Information Verification, and the Demise of Truth,” arXiv, 2026, arxiv.org/abs/2601.11507, accessed 31 January 2026.
[95] Smith, E., et al., “A Brain Capital Grand Strategy: toward economic reimagination,” Molecular Psychiatry, 2020, pmc.ncbi.nlm.nih.gov/articles/PMC8244537, accessed 29 January 2026.
[96] Kulveit, J., et al., “Gradual Disempowerment: Systemic Existential Risks from Incremental AI Development,” arXiv, 2025, arxiv.org/abs/2501.16946, accessed 8 March 2026.
[97] Chesterman, S., “Lawful but Awful: Evolving Legislative Responses to Address Online Misinformation, Disinformation, and Mal-Information in the Age of Generative AI,” The American Journal of Comparative Law, 2024, academic.oup.com/ajcl/article/72/4/933/8189554, accessed 23 February 2026.
[98] DSA Civil Society Coordination Group, “DSA Civil Society Coordination Group Publishes an Initial Analysis of the Major Online Platforms’ Risks Analysis Reports,” Center for Democracy & Technology, 2025, cdt.org/insights/dsa-civil-society-coordination-group-publishes-an-initial-analysis-of-the-major-online-platforms-risks-analysis-reports, accessed 9 March 2026.
[99] For structural critiques of risk-based platform governance as a regulatory design choice, arguing that it inherently delegates political discretion to regulated entities, see Rachel Griffin, ‘Governing platforms through corporate risk management: the politics of systemic risk in the Digital Services Act’, European Law Open, 4(2), 2025, 223–253; Andrea Palumbo, ‘Systemic Risk Management and the Constitutional Limits of Delegating Political Discretion: An Analysis of the DSA and the AI Act’, European Journal of Risk Regulation, 2025, 1–22.
[100] Panoptykon Foundation, “DSA vs. Reality: Are children safer online?,” Panoptykon Foundation, 2026, en.panoptykon.org/dsa-vs-reality-are-children-safer-online-ep-hearing, accessed 18 March 2026.
[101] Haimson, O.L., et al., “Disproportionate Removals and Differing Content Moderation Experiences for Conservative, Transgender, and Black Social Media Users: Marginalization and Moderation Gray Areas,” Proceedings of the ACM on Human-Computer Interaction, 2021, oliverhaimson.com/PDFs/HaimsonDisproportionateRemovals.pdf, accessed 6 February 2026.
[102] Díaz, Á., et al., “Double Standards in Social Media Content Moderation,” Brennan Center for Justice, 2021, brennancenter.org/our-work/research-reports/double-standards-social-media-content-moderation, accessed 14 March 2026.
[103] Gomez, J.F., et al., “Algorithmic Arbitrariness in Content Moderation,” arXiv, 2024, arxiv.org/abs/2402.16979, accessed 7 February 2026.
[104] BBC News, “Molly Russell: Instagram posts seen by teen were safe, Meta says,” BBC News, 2022, bbc.co.uk/news/uk-63034300, accessed 1 March 2026.
[105] Gibson, K., “Meta to end fact-checking on Facebook and Instagram, replacing it with community-driven system akin to Elon Musk’s X,” CBS News, 2025, cbsnews.com/news/meta-facebook-instagram-fact-checking-mark-zuckerberg, accessed 6 March 2026.
[106] Albergotti, R., “Meta’s Fact-Checking Change Could Lead to More Misinformation on Facebook and Instagram,” Time, 2025, time.com/7205332/meta-fact-checking-community-notes, accessed 12 February 2026.
[107]Ecker, U.K.H., et al., “The psychological drivers of misinformation belief and its resistance to correction,” Nature Reviews Psychology, 2022, nature.com/articles/s44159-021-00006-y, accessed 20 February 2026.
[108] Furizal, et al., “Social, legal, and ethical implications of AI-Generated deepfake pornography on digital platforms: A systematic literature review,” Social Sciences & Humanities Open, 2025, sciencedirect.com/science/article/pii/S2590291125006102, accessed 5 February 2026.