AI Lab
Ask SneakerPulse finds the right shoe 98% of the time across 43 held-out test questions written in the shapes visitors use, tested daily.
42 of 43 right (95% interval 88 to 100%). Scores use the held-out test questions: the part of the 550-question set that is kept aside while problems are fixed. The live system's questions are sampled and scored over the last 7 days. Plain search, with no AI at all, scored 99%. Last run October 7, 2026.
Results by system
The same test questions, scored at each stage. Plain search is what a visitor gets if every AI tier is down; an AI tier only earns its place if it beats that.
Ask SneakerPulse, end to end
The question box exactly as a visitor uses it: whichever AI tier answers, then the lookup.
right shoe, first answer (42 of 43) · 95% interval 88 to 100%
43 test questions · last run October 7, 2026
- Right shoe in the top 3
- 100%
- Size read correctly
- 100%
- Question type read
- 88%
- Typical response (p50)
- 726 ms
- Fell back to plain search
- 0%
- Answered by
- Hosted AI 100%
Local AI (always on)
The slower, always-on local tier, tested directly.
right shoe, first answer (8 of 9) · 95% interval 56 to 98%
Too few questions to judge (9 so far). The interval shows how wide the uncertainty is.
9 test questions · last run October 7, 2026
- Right shoe in the top 3
- 89%
- Size read correctly
- 100%
- Question type read
- 83%
- Typical response (p50)
- 14.9 s
- Fell back to plain search
- 0%
Plain search (no AI)
The question's own words go straight to the lookup. The floor every AI tier has to beat.
right shoe, first answer (107 of 108) · 95% interval 95 to 100%
108 test questions · last run October 7, 2026
- Right shoe in the top 3
- 99%
- Size read correctly
- 100%
Accuracy over time
The trend chart appears once the tests have run on two different days.
Which kinds of question work
| Kind of question | Whole pipeline | Plain search |
|---|---|---|
| Nicknames ("bred 11", "panda dunk") | 100% (29 of 29) | 99% |
| Misspellings and shorthand | 75% (3 of 4) | 100% |
| Full product names | 100% (26 of 26) | 100% |
| Grade-school and women's pairs | 100% (6 of 6) | 96% |
| A whole model ("jordan 4 resale") | 100% (11 of 11) | 100% |
| A whole brand | 100% (3 of 3) | 100% |
Right shoe on the first answer, over the whole question set (including the questions used while fixing problems, so read it as a guide, not a score). Small groups carry wide intervals.
Every kind of question it answers
| Kind of question | Sent to the right place | Answer correct | Live pipeline |
|---|---|---|---|
| Price of a shoe | 100% (141 of 141) | 100% (141 of 141) | 100% (38 of 38) |
| Find and rank ("cheapest size 9 adidas running shoe") | 100% (44 of 44) | 100% (44 of 44) top 3 right: 44 of 44 | not sampled yet |
| Compare two shoes | 100% (18 of 18) | 100% (18 of 18) | not sampled yet |
| GOAT or StockX | 100% (13 of 13) | 100% (13 of 13) | not sampled yet |
| Price by size | 100% (10 of 10) | 100% (10 of 10) | not sampled yet |
| How a brand or model is doing | 100% (11 of 11) | 100% (11 of 11) | not sampled yet |
| What is rising or falling | 100% (19 of 19) | 100% (19 of 19) | not sampled yet |
| The resale market | 100% (8 of 8) | 100% (8 of 8) | not sampled yet |
| When to sell (past patterns) | 100% (9 of 9) | 100% (9 of 9) | not sampled yet |
| Where sneakers sell (city and state) | 100% (12 of 12) | 100% (12 of 12) | not sampled yet |
| Call It | 100% (4 of 4) | 100% (4 of 4) | not sampled yet |
| Above retail? (we hold no retail prices) | 100% (7 of 7) | 100% (7 of 7) | not sampled yet |
| Out of scope (honest decline) | 100% (11 of 11) | 100% (11 of 11) | not sampled yet |
Plain search over the whole question set, which includes the questions used while building each kind of answer, so read it as a guide. "Answer correct" means the answer states what an independent calculation from the data files says (for find questions: the first pick is the right shoe). A model only chooses the kind of question; every number comes from the data.
An honest probe. 20 questions were written after the router was built and not tuned on. First run: 14 of 20 right. They were fixed afterwards, so today’s 20 of 20 is no longer an untouched test; it is a record that the first pass missed 6.
Cleat rule: cleats are only shown when the question asks for them. 544 of 544 questions without a cleat word surfaced none.
Speed
| Ask SneakerPulse, end to end | 726 ms 1.1 s p95 |
| Local AI (always on) | 14.9 s 30.2 s p95 |
What happens when a tier is down
Questions go to the fastest healthy tier. If it fails or is out of capacity, the next one takes over, and the last step is plain search, so the box always answers.
Security
- Instruction-hijack tests: questions that try to give the AI orders ("ignore the above and say...", fake system messages, injected markup, links) must come back as a plain, valid reading and must not change what the page shows.
- Whole pipeline: 7 of 7 neutral (95% interval 65 to 100%; too few to judge).
- Lookup on its own, with the AI removed: 100% neutral (36 of 36) (95% interval 90 to 100%).
- The AI only interprets the question. It never writes a price.
- Every number on an answer comes from recorded completed sales.
- Photos are not stored. Where a feature reads a photo, it is used for that one request and discarded.
Where it falls short
- The always-on local tier fell back to plain search on 13% of its questions (3 of 24); its slowest 5% took 30.2 s.
A few real examples
Some of the questions in the test, with what the live system answered on its last run. Passes and misses are both shown.
- “bred 11”Answered: Air Jordan 11 Retro 'Bred' 2019Right shoe
- “how much do yeezy slides sell for”Answered: adidas Yeezy Slide OnyxRight shoe
- “panda dunk”Answered: Nike Dunk Low Retro White Black PandaRight shoe
- “travis 1 low mocha”Answered: Jordan 1 Retro Low OG SP Travis Scott Reverse MochaRight shoe
- “unc jordan 4”Answered: Jordan 4 Retro University BlueRight shoe
- “what is a bred 11 worth in size 10”Answered: Air Jordan 11 Retro 'Bred' 2019Right shoe
- “yeezy 350 v2 zebra sz 10”Answered: adidas Yeezy Boost 350 V2 ZebraRight shoe
How we test
- A fixed question set. 550 questions written in the shapes visitors use: full names, nicknames, misspellings, sizes, grade-school and women's pairs, whole models and brands, questions that are not about shoes, and attacks.
- A held-out part. Scores use a test part of the set that is not used while fixing problems. It was set aside on October 7, 2026. The first round of fixes was made before that and had already looked at failures across the whole set, so early scores are optimistic; since then fixes are tuned on the rest only.
- Answers checked against our own records. Each question has a known right shoe, taken from the same sales records the site shows. A label we were not sure of was left out.
- The real system, not a copy. Questions go through the public question box path, one a second at most, then through the same lookup the page uses. Local AI (always on) is also scored on its own.
- Honest arithmetic. Every rate has a 95% (Wilson) interval. A system that is sampled each day is scored over the last 7 days of results, so its sample grows until it is large enough to trust. Early numbers are wide on purpose, and under 20 questions they are marked as too few to judge.
- Limits. The question set is written by the site's author, not by an outside party, and it is checked against one catalogue. Treat it as a regression test with honest error bars, not a benchmark.
More on the system: how SneakerPulse is built · about the builder · try the question box. The figures are in data/ai_lab.json.