Human-Agent Observatory Methodology

Direct answer

The Human-Agent Observatory produces adjudicated claims, not a stream of chat logs. An experiment asks one question. An episode is one bounded run. Independence means later runs are not the same thread wearing a hat. Replication state records whether a later independent run held, held only in direction, or failed. Production pages are written closed-world from a locked packet—the same fail-closed habit as the rest of Datapiphany’s authority work.

This page explains that model. It does not publish prompts, scoring code, or raw episode JSON. The September 2026 public lock is 48 running-shoe episodes.

Where this sits in Datapiphany’s method

Human-Agent work does not replace Signal → Relationship → Meaning → Opportunity → Decision. It asks what happens when the signal path itself is mediated: the agent’s retrieval and ranking become part of the cultural and commercial system under study.

If the agent’s output is treated as “what consumers want,” you have smuggled a recommendation set into a demand claim. That is the same class of error as treating conversation volume as demand (Cultural Signal vs Demand).

Units

Experiment. A research question, domain, object, surfaces in scope, what is varied, what is held constant, and what would count as an outcome. One experiment may contain many episodes. Locked IDs: HAD-EXP-001 through HAD-EXP-005.

Episode. A single run with enough context to be audited: date, surface, model/product state if known, account state if it matters, and the outcome measures. An episode is not a finding. Public N on this lock: 48.

Finding. A bounded claim with status, limits, and pointers to the experiment(s) that support it. Findings can be independently retrieved. They should not require the whole Lab to be intelligible.

Outcome measure

The primary allocation outcome in the locked running-shoe cells is first-choice product/brand.

That is why the public 63% figure, the 4/4 pair counts, and the cross-surface map can be read together without treating any cell as market share.

Independence

Independence concerns contamination between runs or conversations. A changed contextual constraint is an experimental treatment; it does not by itself make a run independent. Recutting one conversation into eight “tests” is not independence.

HAD-EXP-004 is the concrete case: four functional baselines were run in separate Plus Chat threads rather than sharing one conversation state. They did not all agree. HAD-EXP-002 does not add episodes; it scores cross-agent disagreement on HAD-EXP-001 cells.

Surfaces and product names

Claims about a platform’s present behavior should name:

On this lock:

If a dedicated mode was not observed in the tested interface, the public sentence is that it was not observed there—not that the product does not exist, not that it was tested and failed, and not that it can be backfilled from product cards in ordinary chat.

Evidence classes (public vocabulary)

Aligned to Datapiphany’s existing refusal to flatten confidence:

ClassMeaning
ObservedDirectly supported by the locked episode/experiment
PartialDirection holds; not fully (PARTIALLY_REPLICATES is this class for Finding #001 on consumer surfaces)
ReplicatedObserved again under sufficiently independent conditions
InferredInterpretation beyond the literal observation
HypothesisTestable, not established
ForecastFuture-oriented; must not be written as present fact
UnknownNot resolved
Insufficient dataThe designed test could not be run (Finding #004)
RejectedEvidence does not support it

Public prose should keep these legible without reading like a lab notebook.

Adjudication and lock

Research-time may use the open web. After a proof package is locked, production is closed-world: no new evidence IDs, no quietly promoted claims, no rehabilitating a rejected edge because the copy would land better.

This page is based on a locked September 2026 evidence snapshot. Material evidence changes reopen the research package rather than silently rewriting the claim.

What this methodology refuses to disclose

Proprietary prompts, raw transcripts, internal scores, account chrome, and client applications. Credibility comes from bounded public claims, not from dumping the Lab.

Discuss Human → Agent implications for your category