kpopy demon hunter - Shopping Copilot Plan

Simple slide-style plan for the TikTok TechJam Track 4 backend agent: what it does, why it works, what we compared, and how teammates run it.

No paid APIs Offline default Local UI for demo Windows + macOS/Linux scripts TechnicalScore 0.955300

1. What The Challenge Wants

Challenge requirementWhat that means in normal wordsOur answer
Conversational searchThe customer can be vague and speak over multiple turns.Keep session memory and ask useful questions.
Buying vs browsingSome users know exact constraints; others are exploring.Narrow hard constraints but keep broad fallback candidates.
Intent overrideThe customer may change their mind.Replace old constraints instead of mixing them with new ones.
Backend-only scoringUI is not the judged product.Focus on the official Python agent API and use UI only for demo video.

2. Scoring Cheat Sheet

HitRate@10Did we find it?

Correct product appears somewhere in our 10 returned IDs.

MRRHow high?

Rank 1 is best. Rank 10 is weaker but still useful.

MTTCHow fast?

Average turns before the evaluator sees the correct product.

TechnicalScore = 0.50 * HitRate@10 + 0.30 * MRR + 0.20 * Efficiency

3. Method Comparison

MethodWhat it meansPaid API?Expected scoreUse?
Starter BM25Simple keyword search.No0.1067No
Category + memoryRemember previous turns and use category clues.Noabout 0.25Foundation only
Ask every turnAlways ask a valid attribute question while recommending.Noabout 0.69Yes, core behavior
Fast exact + lexicalUse category, exact constraints, words, and fallback ranking.No0.852704Previous default
Fast + confidence gateAdd semicolon-safe constraints, top-50 reranking, position matching, and wait for enough evidence before submitting a scored slate.No0.955300Submit
Dense embeddingsOptional semantic search model.NoMust beat default firstResearch only
LightGBM rerankerOptional trained ranking model.NoMust beat default firstResearch only
Hosted LLMExternal model call for ranking or rewriting.Likely yesNot neededAvoid

4. What Are We Actually Using?

QuestionAnswer
Are we using an LLM?No hosted LLM is used in the submitted default. No paid API call is required.
Are we using BM25?BM25 means keyword search. We tested BM25-style/hybrid code, but the default is not plain BM25.
What is ranked retrieval?Give each candidate product a score, sort by that score, and return the best 10 IDs.
What is the main search method?Category filtering, exact disclosed constraints, lexical/token overlap, and popularity fallback.
Are we training a model?Not for the submitted default. LightGBM training is a research branch only.
Why this method?It is faster, cheaper, reproducible, and currently beats the heavier research path.

5. Marketplace Research

PlatformWhat they doLesson for us
TaobaoQwen/Taobao-style conversational shopping and follow-up questions.Shopping should be a guided conversation, not a single keyword box.
LazadaLazzieChat/AI Lazzie gives product suggestions from natural questions.Keep suggestions grounded in product facts.
ShopeeSearch/recommendation systems and conversational discovery integrations.Embeddings and recommendation signals are worth researching offline.
AmazonRufus/Alexa uses query understanding, retrieval, product facts, and ranking.Use a funnel: parse, retrieve from multiple routes, rank Top 10.
Detailed branch docs: docs/research/marketplace_search_patterns.md, docs/research/model_avenues.md, docs/research/cross_project_eval_lessons.md, and docs/evaluation_runbook.md.

6. 95% TechnicalScore Reality Check

MetricCurrent clean resultRough target for 95% TechnicalScore
HitRate@101.000about 1.000
MRR0.937about 0.930+
MTTC2.290about 2-3 turns is acceptable if rank 1 improves
30-session research variantHitRate@10MRRMTTCTechnicalScoreDecision
Hybrid full0.8000.6898153.9333330.748278Below default
Hybrid no_dense0.7666670.6814814.4666670.718444Below default
Hybrid lexical_only0.1000.04142910.0000000.082429Reject
Fast default before tie-breaks0.9600.6813472.5850000.852704Previous best
Fast rank tie-break research1.0000.7291071.5250000.908232Previous 90% research result
Fast confidence gate1.0000.9370002.2900000.955300Current default on main
Latest main result is now 0.955300 TechnicalScore. HitRate stayed at 1.0; MRR jumped because the agent waits for more evidence before submitting recommendations.

7. Architecture Diagram

Submitted Offline System

customer message
  |
  v
official evaluator / demo UI
  |
  v
starter.agent.Agent
  |
  v
FastAgent
  |
  +-- parse message
  +-- remember session state
  +-- ask one useful attribute
  |
  v
category + exact + lexical + popularity
  |
  v
rank products
  |
  v
if enough evidence: 10 product IDs + next question
else: ask one more question first

Experiment Decision Flow

new idea
  |
  v
run tests + evaluator
  |
  v
compare score table
  |
  +-- score > 0.955300 and stable -> consider submit
  |
  +-- score <= 0.955300 or unstable -> keep as research
  |
  v
never use paid API by default

8. Current Result Table

MetricValueEasy reading
HitRate@101.000000The agent includes the hidden product in Top 10 for every public session.
MRR0.937000The hidden product is usually rank 1 after enough details are known.
MTTC2.290000The agent waits a little longer so the first scored slate is stronger.
TechnicalScore0.955300Above 0.95 and above our 0.80 acceptance target.
Token usage0No LLM API calls are used in the submitted path.

9. File Map

File/folderPurposeWhy teammates care
starter/agent.pyOfficial import path.This must stay valid for submission.
agent/fast_agent.pyDefault scoring agent.Main logic for the submitted method.
agent/state.pyConversation memory.Prevents forgetting or mixing old intent.
agent/question.pyClarification policy.Controls what the agent asks next.
tools/eval.pyScore reporting.Use it to compare methods.
scripts/*.ps1Windows commands.For Bryan / Windows machines.
scripts/*.shmacOS/Linux commands.For teammates on other machines.
ui/Demo UI.Use for the video, not for scoring.

10. Team Setup Commands

Windows

git pull origin main
python -m pip install -r requirements.txt
.\scripts\setup_local_data.ps1 -DownloadOfficial
.\scripts\verify_submission.ps1 -WithData
.\scripts\demo.ps1 -Fixture

Fixture note: if the UI says Demo fixture catalog: 13 products ready, that is expected. The fixture is only for fast video recording.

macOS / Linux

git pull origin main
python -m pip install -r requirements.txt
sh scripts/setup_local_data.sh --download-official
sh scripts/verify_submission.sh --with-data
sh scripts/demo.sh --fixture

Official scoring: use scripts/evaluate.* with the downloaded 50,000-product catalog. The 13-product UI is not the scored backend run.

11. Useful vs Not Useful Right Now

UsefulWhy
Keep no-paid-API default.Simple, cheap, reproducible, no private credentials.
Show comparison table in Devpost/video.Proves we tested options and chose the best measured method.
Use UI for video only.Looks good in demo while keeping backend scoring clean.
Use `.ps1` and `.sh` script pairs.Teammates can run the project on Windows, macOS, or Linux.
Not usefulWhy
Paid LLM APIs.Cost, credential, and network risk; not needed for current score.
Unmeasured model changes.Could reduce score or break deadline safety.
Overbuilding UI.The challenge is evaluated through backend APIs.

12. External Repo Probe Summary

Subagents inspected three public Track 4-style repositories in temporary clones. We used them only as a sanity check for documentation and risk posture; we did not copy code or use their scores as evidence.

RepoUseful lessonDecision for us
naijovan/Star-Labs-TechJam_TikTok_Track-4Strong no-paid-API disclosure, score table, fallback notes.Keep our own measured 0.955300 result and document optional model paths as research only.
rheaaas11/TikTok-TechJam-HackathonClear local retrieval, contract checks, and scenario metric reporting.Keep README/evaluator commands obvious for teammates and judges.
KRAZYZECTRON/Tiktok-JamOptional dense/local-LLM experiments were disabled by default.Do not ship dense or LLM routes unless they beat the offline FastAgent.

13. Demo Video Script

ShotShowSay
1README topWe are team kpopy demon hunter, building a backend shopping copilot.
2Method tableWe tested the starter, memory, questioning, dense research, and LTR research.
3Architecture diagramThe submitted path is offline ranked retrieval, not a paid LLM API.
4Evaluator tableThis is the official-style 50,000-product score.
5Local UIThe UI fixture has 13 products for recording; the evaluator uses the real catalog.
6No API sectionZero token usage, no hosted model, and teammates can run it with ps1 or sh scripts.

14. Remaining Tasks

TaskOwnerStatus
Confirm no open final PR exists before final packaging from main.Team / @yinasaurusNo open PRs at last check
Confirm individual member names and contribution split.TeamFilled in README and Devpost draft
Record and upload public YouTube video.Demo ownerStill needed
Paste Devpost draft and links.SubmitterStill needed
Build final zip plus .zip.sha256 checksum.EngineeringRun package script from latest main
Keep score above 0.80 and avoid paid APIs.EngineeringCurrently satisfied at 0.955300
Deadline: submit on Devpost before 2026-09-01 12:00 Singapore time.