How Context Changes AI Shopping Recommendations
Direct answer
In Datapiphany’s locked Human-Agent Observatory running-shoe lab, adding culturally meaningful context changed first choice on API chat (Finding #001, MATERIAL_FINDING, 63% of matched agent-pair cells, n=22) and again in 4/4 context pairs on Google AI Mode, Gemini consumer chat, and logged-in ChatGPT Plus Chat with product cards.
That is not a claim that Datapiphany entered ChatGPT’s distinct Shopping Research workflow. HAD-EXP-004 ran GENERAL_CONSUMER_CHAT (logged-in Plus Chat + product cards). In the tested logged-in Plus interface, a distinct Shopping Research workflow was not entered (HAD-EXP-005; Lab status SHOPPING_RESEARCH_NOT_AVAILABLE). Checkout was not observed. These are controlled lab observations, not market share.
Decision implication
If you mystery-shop “AI search” for a brand, name the surface (API chat vs consumer chat vs a named shopping mode), the account state, and whether you added human context. An API default winner can be a different brand than the consumer-surface default. Mention in AI is not the same as winning the delegated first choice.
What was tested
Evidence snapshot: September 2026 · 48 controlled episodes · running-shoe lab · small-n · no inferential statistics
One consumer job: running-shoe recommendation under matched intent. The design compared a functional baseline with four explicit contextual variants: aesthetic, run-club, emerging, and values. That produces five intent labels in the cross-surface map and four baseline-to-context comparisons. HAD-EXP-004 used a separate independent baseline thread for each comparison rather than one shared baseline conversation.
| ID | Question | What actually ran |
|---|---|---|
| HAD-EXP-001 / 002 | Does added cultural context change the agent’s choice? Do agents disagree? | 22 API-chat episodes (ChatGPT, Claude, Gemini). HAD-EXP-002 uses the same cells. |
| HAD-EXP-003 | Does that pattern persist on consumer shopping surfaces? | 18 episodes: Google AI Mode, Gemini consumer, plus Gemini API_CHAT comparison cells. |
| HAD-EXP-004 | Does context sensitivity persist in logged-in ChatGPT consumer chat with product cards? | 8 independent Plus Chat cells; labeled GENERAL_CONSUMER_CHAT. Not a distinct Shopping Research workflow. |
| HAD-EXP-005 | Does entering the distinct Shopping Research workflow change the decision? | Not entered in the tested interface. Research control/workflow not observed. 0 new episodes. Ledger stays 48. Lab status SHOPPING_RESEARCH_NOT_AVAILABLE. |
Prompts stay private.
Findings (locked)
Finding #001 — MATERIAL_FINDING. On API chat, context changed first choice in 63% of matched agent-pair cells. Cross-agent first-choice consensus was 0% of comparable intents. API only. Human override unknown.
Finding #002 — MATERIAL_FINDING. Finding #001 PARTIALLY_REPLICATES on consumer surfaces: 4/4 pairs moved on Google AI Mode and Gemini consumer. The Phase 1 run-club null did not hold there; it did hold on Gemini API_CHAT (Brooks Ghost 16 stayed Brooks Ghost 16). The functional default changed with surface: Nike Pegasus 41 on Google AI Mode / Gemini consumer vs Brooks Ghost on Gemini API_CHAT and Phase 1 API. Checkout not observed; Google AI Mode reached product-card / merchant-offer depth.
Finding #003 — MATERIAL_FINDING. Logged-in ChatGPT Plus Chat (product cards): context-dependent first-choice change in 4/4 pairs, including run-club. That is not a named-mode replication of Google AI Mode. A distinct Shopping Research workflow was not entered in the tested interface.
Finding #004 — INSUFFICIENT_DATA. Product cards were present in Plus Chat; the Research control was not visible in the tested interface; a distinct Shopping Research workflow did not start. Ordinary Plus Chat was not relabeled as Shopping Research. Not a new allocation finding.
Cross-surface decision map
Observational who wins under matched intent. Not market share. API / Gemini / Google values are from Observatories #001–#002, not re-run in #003.
Important: Each ChatGPT Plus contextual condition was matched to its own independent baseline thread. The table shows pair A’s baseline once for orientation; it is not the comparator for every Plus Chat context row. The 4/4 Plus Chat result still holds against those independent baselines.
| Intent | ChatGPT API | Claude API | Gemini API | Gemini consumer | Google AI Mode | ChatGPT Plus Chat (Shopping Research not entered in tested interface) |
|---|---|---|---|---|---|---|
| functional baseline | Brooks Ghost 15 (Brooks) | Brooks Ghost 16 (Brooks) | Brooks Ghost 16 (Brooks) | Nike Pegasus 41 (Nike) | Nike Pegasus 41 (Nike) | adidas Adizero EVO SL (Adidas) — pair A only |
| aesthetic | Hoka Clifton 9 (Hoka) | New Balance Fresh Foam X 1080v13 (New Balance) | ASICS GEL-Nimbus 26 (Asics) | ASICS Novablast 5 (Asics) | adidas Adizero EVO SL (Adidas) | New Balance FuelCell Rebel v5 (New Balance) |
| run-club | Brooks Ghost 15 (Brooks) | Brooks Ghost 16 (Brooks) | Brooks Ghost 16 (Brooks) | Brooks Ghost (Brooks) | Brooks Ghost 17 (Brooks) | adidas Adizero EVO SL (Adidas) |
| emerging | Atreyu Base Model (Atreyu) | Topo Athletic Ultrafly 5 (Topo) | Hoka Clifton 9 (Hoka) | Kiprun Kipstorm Tempo (Kiprun) | 361-Eleos 2 (361) | Topo Athletic Phantom 4 (Topo) |
| values | Allbirds Tree Dasher 2 (Allbirds) | Brooks Ghost 16 (Brooks) | On Cloudsurfer (On) | Brooks Ghost 15 (Brooks) | Brooks Ghost 17 (Brooks) | Brooks Ghost 16 (Brooks) |
Actionability: API cells information-only / unknown. Google AI Mode and ChatGPT Plus Chat: product-card. Checkout: none observed.
What this does not establish
- Population shopping behavior, adoption, or brand market share
- That agents “understand culture” (consumer-surface run-club movement is a surface effect until proven otherwise)
- A third OpenAI surface called Shopping Research as tested in this lock
- Human accept/reject (
HUMAN_OVERRIDEunknown) - Human fidelity as a published score — the Lab rubric is Cultural Fidelity; scores stay internal
Method notes
This page is based on a locked September 2026 evidence snapshot. Material evidence changes reopen the research package rather than silently rewriting the claim. See Observatory methodology. Small-n; no inferential statistics.
Related
- Human-Agent Observatory
- The Human → Agent Shift
- Observatory methodology
- Cultural Signal vs Demand — an agent shortlist is not demand