ACE Journal

Labeler Welfare and Exploitation in RLHF Data Supply Chains

Abstract

Reinforcement learning from human feedback (RLHF) has become the dominant technique for aligning large language models with human preferences. It depends on a global workforce of data labelers who rate model outputs, identify harmful content, and provide comparative preference judgments at scale. This workforce is largely invisible to the public discourse around AI development, geographically concentrated in lower-income countries, often employed through layered subcontracting arrangements that dilute accountability, and routinely exposed to psychologically harmful content without adequate support. The ethics of RLHF as an alignment technique cannot be separated from the labor conditions under which its training signal is produced. This article examines what is known about labeler working conditions, the structural features that produce exploitation, and the governance levers that could address it.

The Structure of the RLHF Labor Market

Large AI labs - including OpenAI, Anthropic, Google DeepMind, and Meta - contract data labeling work to intermediaries including Scale AI, Surge AI, Remotasks, and iMerit, among others. These intermediaries in turn recruit from global crowdwork platforms or operate dedicated facilities, primarily in Kenya, the Philippines, India, Venezuela, and Pakistan. Wage rates, working conditions, and psychological support vary substantially across this chain, and the contracting structure typically means that the AI lab’s code of conduct applies to the direct contractor but is not consistently enforced down the subcontracting stack.

Time Investigation’s 2023 reporting on Sama’s work for OpenAI in Kenya - labeling graphic violence, child exploitation material, and other traumatic content for model safety training - put the psychological costs of this work in public view for the first time. Since then, additional reporting has documented similar conditions at Scale AI facilities and at crowdwork platforms where tasks involving harmful content are distributed without content warnings or support resources. A 2024 peer-reviewed study in the Journal of AI Ethics estimated that between 15 and 30 percent of RLHF labelers are exposed to severely distressing content in a given month, with access to psychological support far lower than that fraction.

Structural Features That Produce Exploitation

Several structural features of the RLHF market create conditions for exploitation independent of the intentions of individual companies. Piece-rate compensation tied to throughput creates incentives to process tasks quickly, reducing the time labelers spend on the psychological processing that harmful content requires. Gig-based employment structures exclude labelers from the labor protections that govern employment relationships in most jurisdictions. Geographic concentration in countries where regulatory enforcement of labor standards is weak means that even nominally applicable laws rarely constrain working conditions. And the opacity of subcontracting chains means that brands with published ethical commitments can remain insulated from the conditions their supply chains produce.

The power asymmetry is compounded by non-disclosure agreements that labelers are typically required to sign. Workers who experience harm from content exposure or unfair termination have limited legal recourse and strong contractual disincentives to speak publicly.

Governance Levers and Their Limits

Several interventions could improve labeler welfare without eliminating the RLHF supply chain. Extended duty-of-care obligations - requiring AI labs to enforce labor standards throughout their subcontracting chains, with auditable certification - would be the most direct mechanism, analogous to supply chain due diligence requirements that apply to other industries in the EU under the Corporate Sustainability Due Diligence Directive (CSDDD). Mandatory content warning and psychological support protocols, with minimum standards set by regulation rather than left to contractor discretion, would address the acute harm of traumatic content exposure. Wage floor requirements pegged to purchasing power parity benchmarks would partially correct the geographic arbitrage that keeps labeler compensation low.

Progress on any of these has been slow. The EU AI Act’s supply chain provisions focus on data provenance and technical documentation, not on labor conditions of the humans who produce training signal. No major AI lab has made its labeler welfare standards auditable by independent third parties. The burden of demonstrating that RLHF-based alignment is compatible with ethical labor practices still rests almost entirely on the advocacy of researchers and journalists rather than on structural governance.