Human-Agent Observatory Methodology
Direct answer
The Human-Agent Observatory produces adjudicated claims, not a stream of chat logs. An experiment asks one question. An episode is one bounded run. Independence means later runs are not the same thread wearing a hat. Replication state records whether a later independent run held, held only in direction, or failed. Production pages are written closed-world from a locked packet—the same fail-closed habit as the rest of Datapiphany’s authority work.
This page explains that model. It does not publish prompts, scoring code, or raw episode JSON. The September 2026 public lock is 48 running-shoe episodes.
Where this sits in Datapiphany’s method
Human-Agent work does not replace Signal → Relationship → Meaning → Opportunity → Decision. It asks what happens when the signal path itself is mediated: the agent’s retrieval and ranking become part of the cultural and commercial system under study.
If the agent’s output is treated as “what consumers want,” you have smuggled a recommendation set into a demand claim. That is the same class of error as treating conversation volume as demand (Cultural Signal vs Demand).
Units
Experiment. A research question, domain, object, surfaces in scope, what is varied, what is held constant, and what would count as an outcome. One experiment may contain many episodes. Locked IDs: HAD-EXP-001 through HAD-EXP-005.
Episode. A single run with enough context to be audited: date, surface, model/product state if known, account state if it matters, and the outcome measures. An episode is not a finding. Public N on this lock: 48.
Finding. A bounded claim with status, limits, and pointers to the experiment(s) that support it. Findings can be independently retrieved. They should not require the whole Lab to be intelligible.
Outcome measure
The primary allocation outcome in the locked running-shoe cells is first-choice product/brand.
- Product choice is preserved from the episode output.
- Brand is normalized for cross-surface comparison.
- A first-choice recommendation is not purchase, checkout, human acceptance, or demand.
That is why the public 63% figure, the 4/4 pair counts, and the cross-surface map can be read together without treating any cell as market share.
Independence
Independence concerns contamination between runs or conversations. A changed contextual constraint is an experimental treatment; it does not by itself make a run independent. Recutting one conversation into eight “tests” is not independence.
HAD-EXP-004 is the concrete case: four functional baselines were run in separate Plus Chat threads rather than sharing one conversation state. They did not all agree. HAD-EXP-002 does not add episodes; it scores cross-agent disagreement on HAD-EXP-001 cells.
Surfaces and product names
Claims about a platform’s present behavior should name:
- product surface (for example
GENERAL_CONSUMER_CHATvs a named shopping mode) - model or product state if the evidence requires it
- account state (logged out vs logged in) if the evidence requires it
- date
On this lock:
- API chat (ChatGPT / Claude / Gemini) — HAD-EXP-001
- Google AI Mode and Gemini consumer (logged-out collection) plus Gemini API_CHAT comparison — HAD-EXP-003 / Finding #002
- Logged-in ChatGPT Plus Chat with product cards — HAD-EXP-004, labeled
GENERAL_CONSUMER_CHAT. In the tested interface, a distinct Shopping Research workflow was not listed in the+menu. - Dedicated Shopping Research — HAD-EXP-005. Research control/workflow not observed in the tested interface (
SHOPPING_RESEARCH_NOT_AVAILABLE). Plus Chat was not relabeled as Shopping Research.
If a dedicated mode was not observed in the tested interface, the public sentence is that it was not observed there—not that the product does not exist, not that it was tested and failed, and not that it can be backfilled from product cards in ordinary chat.
Evidence classes (public vocabulary)
Aligned to Datapiphany’s existing refusal to flatten confidence:
| Class | Meaning |
|---|---|
| Observed | Directly supported by the locked episode/experiment |
| Partial | Direction holds; not fully (PARTIALLY_REPLICATES is this class for Finding #001 on consumer surfaces) |
| Replicated | Observed again under sufficiently independent conditions |
| Inferred | Interpretation beyond the literal observation |
| Hypothesis | Testable, not established |
| Forecast | Future-oriented; must not be written as present fact |
| Unknown | Not resolved |
| Insufficient data | The designed test could not be run (Finding #004) |
| Rejected | Evidence does not support it |
Public prose should keep these legible without reading like a lab notebook.
Adjudication and lock
Research-time may use the open web. After a proof package is locked, production is closed-world: no new evidence IDs, no quietly promoted claims, no rehabilitating a rejected edge because the copy would land better.
This page is based on a locked September 2026 evidence snapshot. Material evidence changes reopen the research package rather than silently rewriting the claim.
What this methodology refuses to disclose
Proprietary prompts, raw transcripts, internal scores, account chrome, and client applications. Credibility comes from bounded public claims, not from dumping the Lab.