AI Lab

Ask SneakerPulse finds the right shoe 98% of the time across 43 held-out test questions written in the shapes visitors use, tested daily.

42 of 43 right (95% interval 88 to 100%). Scores use the held-out test questions: the part of the 550-question set that is kept aside while problems are fixed. The live system's questions are sampled and scored over the last 7 days. Plain search, with no AI at all, scored 99%. Last run October 7, 2026.

Results by system

The same test questions, scored at each stage. Plain search is what a visitor gets if every AI tier is down; an AI tier only earns its place if it beats that.

Public path

Ask SneakerPulse, end to end

The question box exactly as a visitor uses it: whichever AI tier answers, then the lookup.

98%

right shoe, first answer (42 of 43) · 95% interval 88 to 100%

43 test questions · last run October 7, 2026

Right shoe in the top 3
100%
Size read correctly
100%
Question type read
88%
Typical response (p50)
726 ms
Fell back to plain search
0%
Answered by
Hosted AI 100%

Local AI (always on)

The slower, always-on local tier, tested directly.

89%

right shoe, first answer (8 of 9) · 95% interval 56 to 98%

Too few questions to judge (9 so far). The interval shows how wide the uncertainty is.

9 test questions · last run October 7, 2026

Right shoe in the top 3
89%
Size read correctly
100%
Question type read
83%
Typical response (p50)
14.9 s
Fell back to plain search
0%
The floor

Plain search (no AI)

The question's own words go straight to the lookup. The floor every AI tier has to beat.

99%

right shoe, first answer (107 of 108) · 95% interval 95 to 100%

108 test questions · last run October 7, 2026

Right shoe in the top 3
99%
Size read correctly
100%

Accuracy over time

The trend chart appears once the tests have run on two different days.

Which kinds of question work

Kind of questionWhole pipelinePlain search
Nicknames ("bred 11", "panda dunk")100% (29 of 29)99%
Misspellings and shorthand75% (3 of 4)100%
Full product names100% (26 of 26)100%
Grade-school and women's pairs100% (6 of 6)96%
A whole model ("jordan 4 resale")100% (11 of 11)100%
A whole brand100% (3 of 3)100%

Right shoe on the first answer, over the whole question set (including the questions used while fixing problems, so read it as a guide, not a score). Small groups carry wide intervals.

Every kind of question it answers

Kind of questionSent to the right placeAnswer correctLive pipeline
Price of a shoe100% (141 of 141)100% (141 of 141)100% (38 of 38)
Find and rank ("cheapest size 9 adidas running shoe")100% (44 of 44)100% (44 of 44)
top 3 right: 44 of 44
not sampled yet
Compare two shoes100% (18 of 18)100% (18 of 18)not sampled yet
GOAT or StockX100% (13 of 13)100% (13 of 13)not sampled yet
Price by size100% (10 of 10)100% (10 of 10)not sampled yet
How a brand or model is doing100% (11 of 11)100% (11 of 11)not sampled yet
What is rising or falling100% (19 of 19)100% (19 of 19)not sampled yet
The resale market100% (8 of 8)100% (8 of 8)not sampled yet
When to sell (past patterns)100% (9 of 9)100% (9 of 9)not sampled yet
Where sneakers sell (city and state)100% (12 of 12)100% (12 of 12)not sampled yet
Call It100% (4 of 4)100% (4 of 4)not sampled yet
Above retail? (we hold no retail prices)100% (7 of 7)100% (7 of 7)not sampled yet
Out of scope (honest decline)100% (11 of 11)100% (11 of 11)not sampled yet

Plain search over the whole question set, which includes the questions used while building each kind of answer, so read it as a guide. "Answer correct" means the answer states what an independent calculation from the data files says (for find questions: the first pick is the right shoe). A model only chooses the kind of question; every number comes from the data.

An honest probe. 20 questions were written after the router was built and not tuned on. First run: 14 of 20 right. They were fixed afterwards, so today’s 20 of 20 is no longer an untouched test; it is a record that the first pass missed 6.

Cleat rule: cleats are only shown when the question asks for them. 544 of 544 questions without a cleat word surfaced none.

Speed

Dark bar: typical (median) response. Light bar: slowest 5% (p95). Plain search has no AI step, so it is not timed.
Ask SneakerPulse, end to end
726 ms
1.1 s p95
Local AI (always on)
14.9 s
30.2 s p95

What happens when a tier is down

Questions go to the fastest healthy tier. If it fails or is out of capacity, the next one takes over, and the last step is plain search, so the box always answers.

Each question tries the first tier that is up; if one fails the next takes over.Local AI (fast)tried firstHosted AInext, if the first is downLocal AI (always on)next, if that is downPlain searchno AI, always worksLocal AI (fast)tried firstHosted AInext, if the first is downLocal AI (always on)next, if that is downPlain searchno AI, always works

Security

  • Instruction-hijack tests: questions that try to give the AI orders ("ignore the above and say...", fake system messages, injected markup, links) must come back as a plain, valid reading and must not change what the page shows.
    • Whole pipeline: 7 of 7 neutral (95% interval 65 to 100%; too few to judge).
    • Lookup on its own, with the AI removed: 100% neutral (36 of 36) (95% interval 90 to 100%).
  • The AI only interprets the question. It never writes a price.
  • Every number on an answer comes from recorded completed sales.
  • Photos are not stored. Where a feature reads a photo, it is used for that one request and discarded.

Where it falls short

  • The always-on local tier fell back to plain search on 13% of its questions (3 of 24); its slowest 5% took 30.2 s.

A few real examples

Some of the questions in the test, with what the live system answered on its last run. Passes and misses are both shown.

How we test

More on the system: how SneakerPulse is built · about the builder · try the question box. The figures are in data/ai_lab.json.