Human-Agent Observatory
Direct answer
The Human-Agent Observatory is Datapiphany’s research program on the Human → Agent shift: how discovery, evaluation, recommendation, and action move toward agent mediation, and when that mediation still fails the human it claims to represent.
It is a structured research program, not a blog column or live dashboard. Current public evidence: 48 controlled running-shoe episodes and four adjudicated findings (#001–#003 MATERIAL_FINDING, #004 INSUFFICIENT_DATA). Raw transcripts stay private. In the tested logged-in Plus interface, a distinct Shopping Research workflow was not entered.
Research mission
Evidence snapshot: September 2026 · 48 controlled episodes · running-shoe lab · small-n · no inferential statistics
Question the Observatory exists to answer, in working form:
When an agent completes a consumer task, under what conditions does the output still fail human fidelity—and what does that change for brands and growth decisions?
Sub-questions now partly answered in lab (not as population proof):
- Does explicitly adding contextual preferences change first choice? Yes, in this lab (Finding #001; partial replication on consumer surfaces).
- Must surfaces, models, and account states stay distinct? Yes — API vs consumer UI changed the functional default winner; Plus Chat ≠ Shopping Research.
- Was a distinct Shopping Research workflow entered in the tested Plus interface? No. The Research control/workflow was not observed (
SHOPPING_RESEARCH_NOT_AVAILABLE). - What evidence is still missing for a stronger “agents will choose” claim? Observed delegated action/checkout and human accept/reject behavior. Neither was measured here.
Evidence model
The Observatory produces bounded objects, not vibe:
| Object | Job |
|---|---|
| Experiment | One research question, design, and scope |
| Episode | One bounded run under stated conditions |
| Finding | A claim with status, limits, and replication state |
| Method version | How adjudication was supposed to work at a date |
Public statuses in use on this lock: MATERIAL_FINDING, PARTIALLY_REPLICATES (replication state of #001 on consumer surfaces), INSUFFICIENT_DATA, not observed.
Independence and replication
- Independence — do not treat the same conversation as eight studies. HAD-EXP-004 used independent Plus Chat threads for the functional baseline.
- Replication state — Finding #001
PARTIALLY_REPLICATESon Google AI Mode, Gemini consumer, and ChatGPT Plus Chat; it does not become a different finding number. - Falsifiers — named on each finding report (example: a later dedicated Shopping Research matrix that did or did not match Google AI Mode).
- Surface precision — Plus Chat with product cards is
GENERAL_CONSUMER_CHAT. Do not write “ChatGPT shopping” for that run.
See methodology.
Current public evidence count
Locked public experiments: HAD-EXP-001/002 (API), HAD-EXP-003 (consumer surfaces), HAD-EXP-004 (GENERAL_CONSUMER_CHAT). HAD-EXP-005 did not ingest cells.
Locked episode ledger: 48.
Locked public findings:
| ID | Status | One-line |
|---|---|---|
| Finding #001 | MATERIAL_FINDING | API chat: context changed first choice in 63% of matched cells; 0% cross-agent first-choice consensus (n=22) |
| Finding #002 | MATERIAL_FINDING | #001 PARTIALLY_REPLICATES on Google AI Mode + Gemini consumer (4/4); surface changed the Nike vs Brooks default |
| Finding #003 | MATERIAL_FINDING | Logged-in Plus Chat product cards: 4/4 including run-club; distinct Shopping Research workflow not entered in the tested interface |
| Finding #004 | INSUFFICIENT_DATA | Distinct Shopping Research workflow not entered (SHOPPING_RESEARCH_NOT_AVAILABLE) |
Detail and the cross-surface who-wins map: How context changes AI shopping recommendations.
Public / private boundary
Public, when locked: validated findings, bounded method, selected evidence, definitions, limits, replication status.
Private: raw transcripts, proprietary prompts, unvalidated hypotheses, scoring internals, client diagnostics, account chrome.
This page is based on a locked September 2026 evidence snapshot. Material evidence changes reopen the research package rather than silently rewriting the claim.
Cadence and updates
Update this page when (a) a finding is locked, (b) a finding’s replication state changes, or (c) the method version changes. Do not rewrite it to look fresh.