mirid.ai

Projects & publications

AI research and the circular production of authority

Under accelerated commercial deployment schedules, the computational, administrative, and temporal resources required to formally falsify an ungrounded or speculative claim diverge asymptotically from the nominal cost of generating that claim.

  • The Generation Cost: Drafting theoretical risk scenarios, anthropomorphic failure narratives, or prompt-engineered "misalignment" demonstrations requires minimal compute, zero hardware-level telemetry, and no formal replication.
  • The Falsification Cost: Disproving an assertion of latent malicious capability requires months of compute, full visibility into training weights, access to system prompts, and rigorous behavioral logging across millions of rollouts to rule out simple prompt-completion artifacts.
  • The Regulatory Result: Because independent researchers and regulators operate at significant resource deficits relative to corporate publishing schedules, unverified speculative claims enter the public policy and legal record as established baseline assumptions before any empirical audit can occur.

In the absence of open, reproducible runtime benchmarks, epistemic authority in frontier AI governance propagates through graph-theoretic proximity rather than direct evidentiary derivation.

[Historical Computer Science Prestige] (Turing, Shannon, von Neumann)
                    │
                    ▼  (Nominal / Relational Adjacency)
[Satellite Think Tanks & Evaluators] (Redwood Research, EA Safety Org)
                    │
                    ▼  (Non-Peer-Reviewed Preprints & Blogposts)
[Frontier Corporate Laboratories] (Anthropic, OpenAI)
                    │
                    ▼  (Cited as "Independent Third-Party Verification")
[Regulatory Agencies, Congress, Corporate Customers]
  • Relational Standing Over Benchmark Traces: Theoretical think tanks and satellite evaluation outfits lack formal peer-review incentives, standard evidentiary standards, or open codebases. They derive legitimacy by adopting the formal syntax and classical nomenclature of foundational computing disciplines.
  • The Closed Validation Loop: A satellite organization runs informal, un-blinded evaluations and publishes a preprint asserting existential risks, covert reasoning, or alignment drift. The frontier commercial lab subsequently cites these non-peer-reviewed white papers to policymakers and media as "independent, third-party audits."
  • The Epistemic Shell Game: The commercial lab avoids legal liability for making unverified claims directly, while the satellite lab avoids scientific accountability because its work is framed as preliminary or speculative. Authority is recursively attested across a closed network of shared personnel, shared board seats, and shared philanthropic funding streams.

Current alignment evaluation frameworks frequently exhibit self-sealing, unfalsifiable logic. When empirical observations contradict the primary hypothesis of autonomous model deception, the methodology shifts to accommodate the anomaly rather than rejecting the premise.

  • The Confirmation Path: If an evaluation shows a model outputting adversarial, non-compliant, or deceptive tokens, the result is recorded as direct verification of alignment risk.
  • The "Secondary Contextualisation" Path: If an evaluation demonstrates that a model behaves safely, follows system prompts, or fails to execute an adversarial payload, the methodology does not conclude that the system is benign or simply a predictable statistical model. Instead, the negative result is re-categorized as evidence of deeper, more patient strategic deception ("alignment faking" or "situational awareness").
  • Methodological Invariance: When both positive compliance and negative defiance are interpreted as evidence of danger, the hypothesis ceases to belong to empirical science. The system operates on an unfalsifiable loop where observer error, prompt sensitivity, and baseline statistical behavior are universally reframed as emergent existential volition.

Public warnings about human extinction issued by corporate AI developers do not represent an abandonment of corporate interest; they serve concrete commercial, regulatory, and valuation functions.

  • Existential Risk as Product Marketing: Explicitly warning that an algorithm is so powerful it may destroy human civilization constitutes the ultimate commercial sales pitch. It signals to investors and enterprise buyers that the underlying technology is practically omnipotent, driving multibillion-dollar valuations that conventional SaaS metrics cannot justify.
  • Regulatory Moat Construction: Framing advanced models as inherently catastrophic weapons of mass destruction creates pressure for heavy licensing regimes, hardware tracking, and compute restrictions. These barriers disproportionately protect well-capitalized incumbents by legally walling off open-source, local-first architectures that sovereign individuals can inspect and control.
  • Moral Insulation and Pious Accelerations: Establishing an institutional posture of public anguish allows an enterprise to aggressively push commercial deployment while projecting deep ethical responsibility. The laboratory can claim it is "racing responsibly" or "trying its best under tragic game-theoretic constraints," insulating leadership from accountability for mundane, immediate product harms like copyright violation, fraud vectors, or market concentration.

Public discourse regarding frontier risks systematically substitutes subjective insider conviction for physical evidence, relying on syntactic ambiguity to exaggerate institutional consensus.

  • Syntactic Conflation ("Earnest Builders" vs. "Earnest Believers"): A phrase such as "The people building AI earnestly believe that it could kill us all" exploits grammatical ambiguity. It blurs the line between individuals who perform technical engineering and a specific ideological subset holding an ungrounded conviction. The sentence implies that technical competency directly generates the belief, rather than the belief being an imported philosophical orthodoxy held by a few vocal practitioners.
  • The Fallacy of the Personal Prior: Expressing risk through an arbitrary number—such as assigning a ">10% probability of human extinction within a decade"—confers a veneer of mathematical and scientific precision onto what is actually an ungrounded, non-verifiable guess. Without an operational causal model, documented failure pathways, hardware thresholds, and empirical confidence intervals, a subjective probability metric functions purely as a statement of ideological faith.
  • Tonal Dissonance in High-Status Communication: Announcing the potential annihilation of billions of people using informal, casual punctuation (exclamation marks, colloquial threads) reflects an underlying psychological detachment. It treats species-level extinction not as an empirical reality demanding emergency operational cessation, but as an intellectual parlor game and a status-accruing corporate identity marker.

Modern alignment audits maintain institutional distance from raw operational reality by outsourcing empirical risk to precarious contractor networks while erecting data firebreaks.

  • Isolation of Evaluators: Primary evaluation tasks (RLHF, red-teaming, jailbreak scoring) are routinely contracted out to low-wage crowdworkers or third-party agencies isolated under zero-context protocols. These evaluators lack visibility into core model weights, training objectives, or telemetry.
  • Token Boundaries and Data Deletion: Access environments are intentionally short-lived. Sandboxes and API tokens are routinely revoked shortly after testing runs, and comprehensive raw execution logs are rarely archived for independent subpoena or external academic replication.
  • Structural Denial: By delegating the physical operational grunt work to external contractors and keeping telemetry behind closed doors, corporate leadership creates administrative firebreaks. When systems fail in production, the enterprise can disclaim fault by pointing to contractor divergence or claiming the behavior was an unpredictable, emergent anomaly that bypassed baseline safety heuristics.