1. What The Challenge Wants
| Challenge requirement | What that means in normal words | Our answer |
|---|---|---|
| Conversational search | The customer can be vague and speak over multiple turns. | Keep session memory and ask useful questions. |
| Buying vs browsing | Some users know exact constraints; others are exploring. | Narrow hard constraints but keep broad fallback candidates. |
| Intent override | The customer may change their mind. | Replace old constraints instead of mixing them with new ones. |
| Backend-only scoring | UI is not the judged product. | Focus on the official Python agent API and use UI only for demo video. |
2. Scoring Cheat Sheet
Correct product appears somewhere in our 10 returned IDs.
Rank 1 is best. Rank 10 is weaker but still useful.
Average turns before the evaluator sees the correct product.
TechnicalScore = 0.50 * HitRate@10 + 0.30 * MRR + 0.20 * Efficiency
3. Method Comparison
| Method | What it means | Paid API? | Expected score | Use? |
|---|---|---|---|---|
| Starter BM25 | Simple keyword search. | No | 0.1067 | No |
| Category + memory | Remember previous turns and use category clues. | No | about 0.25 | Foundation only |
| Ask every turn | Always ask a valid attribute question while recommending. | No | about 0.69 | Yes, core behavior |
| Fast exact + lexical | Use category, exact constraints, words, and fallback ranking. | No | 0.852704 | Previous default |
| Fast + confidence gate | Add semicolon-safe constraints, top-50 reranking, position matching, and wait for enough evidence before submitting a scored slate. | No | 0.955300 | Submit |
| Dense embeddings | Optional semantic search model. | No | Must beat default first | Research only |
| LightGBM reranker | Optional trained ranking model. | No | Must beat default first | Research only |
| Hosted LLM | External model call for ranking or rewriting. | Likely yes | Not needed | Avoid |
4. What Are We Actually Using?
| Question | Answer |
|---|---|
| Are we using an LLM? | No hosted LLM is used in the submitted default. No paid API call is required. |
| Are we using BM25? | BM25 means keyword search. We tested BM25-style/hybrid code, but the default is not plain BM25. |
| What is ranked retrieval? | Give each candidate product a score, sort by that score, and return the best 10 IDs. |
| What is the main search method? | Category filtering, exact disclosed constraints, lexical/token overlap, and popularity fallback. |
| Are we training a model? | Not for the submitted default. LightGBM training is a research branch only. |
| Why this method? | It is faster, cheaper, reproducible, and currently beats the heavier research path. |
5. Marketplace Research
| Platform | What they do | Lesson for us |
|---|---|---|
| Taobao | Qwen/Taobao-style conversational shopping and follow-up questions. | Shopping should be a guided conversation, not a single keyword box. |
| Lazada | LazzieChat/AI Lazzie gives product suggestions from natural questions. | Keep suggestions grounded in product facts. |
| Shopee | Search/recommendation systems and conversational discovery integrations. | Embeddings and recommendation signals are worth researching offline. |
| Amazon | Rufus/Alexa uses query understanding, retrieval, product facts, and ranking. | Use a funnel: parse, retrieve from multiple routes, rank Top 10. |
docs/research/marketplace_search_patterns.md, docs/research/model_avenues.md, docs/research/cross_project_eval_lessons.md, and docs/evaluation_runbook.md.6. 95% TechnicalScore Reality Check
| Metric | Current clean result | Rough target for 95% TechnicalScore |
|---|---|---|
| HitRate@10 | 1.000 | about 1.000 |
| MRR | 0.937 | about 0.930+ |
| MTTC | 2.290 | about 2-3 turns is acceptable if rank 1 improves |
| 30-session research variant | HitRate@10 | MRR | MTTC | TechnicalScore | Decision |
|---|---|---|---|---|---|
Hybrid full | 0.800 | 0.689815 | 3.933333 | 0.748278 | Below default |
Hybrid no_dense | 0.766667 | 0.681481 | 4.466667 | 0.718444 | Below default |
Hybrid lexical_only | 0.100 | 0.041429 | 10.000000 | 0.082429 | Reject |
| Fast default before tie-breaks | 0.960 | 0.681347 | 2.585000 | 0.852704 | Previous best |
| Fast rank tie-break research | 1.000 | 0.729107 | 1.525000 | 0.908232 | Previous 90% research result |
| Fast confidence gate | 1.000 | 0.937000 | 2.290000 | 0.955300 | Current default on main |
main result is now 0.955300 TechnicalScore. HitRate stayed at 1.0; MRR jumped because the agent waits for more evidence before submitting recommendations.7. Architecture Diagram
Submitted Offline System
customer message
|
v
official evaluator / demo UI
|
v
starter.agent.Agent
|
v
FastAgent
|
+-- parse message
+-- remember session state
+-- ask one useful attribute
|
v
category + exact + lexical + popularity
|
v
rank products
|
v
if enough evidence: 10 product IDs + next question
else: ask one more question first
Experiment Decision Flow
new idea
|
v
run tests + evaluator
|
v
compare score table
|
+-- score > 0.955300 and stable -> consider submit
|
+-- score <= 0.955300 or unstable -> keep as research
|
v
never use paid API by default
8. Current Result Table
| Metric | Value | Easy reading |
|---|---|---|
| HitRate@10 | 1.000000 | The agent includes the hidden product in Top 10 for every public session. |
| MRR | 0.937000 | The hidden product is usually rank 1 after enough details are known. |
| MTTC | 2.290000 | The agent waits a little longer so the first scored slate is stronger. |
| TechnicalScore | 0.955300 | Above 0.95 and above our 0.80 acceptance target. |
| Token usage | 0 | No LLM API calls are used in the submitted path. |
9. File Map
| File/folder | Purpose | Why teammates care |
|---|---|---|
starter/agent.py | Official import path. | This must stay valid for submission. |
agent/fast_agent.py | Default scoring agent. | Main logic for the submitted method. |
agent/state.py | Conversation memory. | Prevents forgetting or mixing old intent. |
agent/question.py | Clarification policy. | Controls what the agent asks next. |
tools/eval.py | Score reporting. | Use it to compare methods. |
scripts/*.ps1 | Windows commands. | For Bryan / Windows machines. |
scripts/*.sh | macOS/Linux commands. | For teammates on other machines. |
ui/ | Demo UI. | Use for the video, not for scoring. |
10. Team Setup Commands
Windows
git pull origin main
python -m pip install -r requirements.txt
.\scripts\setup_local_data.ps1 -DownloadOfficial
.\scripts\verify_submission.ps1 -WithData
.\scripts\demo.ps1 -Fixture
Fixture note: if the UI says Demo fixture catalog: 13 products ready, that is expected. The fixture is only for fast video recording.
macOS / Linux
git pull origin main
python -m pip install -r requirements.txt
sh scripts/setup_local_data.sh --download-official
sh scripts/verify_submission.sh --with-data
sh scripts/demo.sh --fixture
Official scoring: use scripts/evaluate.* with the downloaded 50,000-product catalog. The 13-product UI is not the scored backend run.
11. Useful vs Not Useful Right Now
| Useful | Why |
|---|---|
| Keep no-paid-API default. | Simple, cheap, reproducible, no private credentials. |
| Show comparison table in Devpost/video. | Proves we tested options and chose the best measured method. |
| Use UI for video only. | Looks good in demo while keeping backend scoring clean. |
| Use `.ps1` and `.sh` script pairs. | Teammates can run the project on Windows, macOS, or Linux. |
| Not useful | Why |
|---|---|
| Paid LLM APIs. | Cost, credential, and network risk; not needed for current score. |
| Unmeasured model changes. | Could reduce score or break deadline safety. |
| Overbuilding UI. | The challenge is evaluated through backend APIs. |
12. External Repo Probe Summary
Subagents inspected three public Track 4-style repositories in temporary clones. We used them only as a sanity check for documentation and risk posture; we did not copy code or use their scores as evidence.
| Repo | Useful lesson | Decision for us |
|---|---|---|
naijovan/Star-Labs-TechJam_TikTok_Track-4 | Strong no-paid-API disclosure, score table, fallback notes. | Keep our own measured 0.955300 result and document optional model paths as research only. |
rheaaas11/TikTok-TechJam-Hackathon | Clear local retrieval, contract checks, and scenario metric reporting. | Keep README/evaluator commands obvious for teammates and judges. |
KRAZYZECTRON/Tiktok-Jam | Optional dense/local-LLM experiments were disabled by default. | Do not ship dense or LLM routes unless they beat the offline FastAgent. |
13. Demo Video Script
| Shot | Show | Say |
|---|---|---|
| 1 | README top | We are team kpopy demon hunter, building a backend shopping copilot. |
| 2 | Method table | We tested the starter, memory, questioning, dense research, and LTR research. |
| 3 | Architecture diagram | The submitted path is offline ranked retrieval, not a paid LLM API. |
| 4 | Evaluator table | This is the official-style 50,000-product score. |
| 5 | Local UI | The UI fixture has 13 products for recording; the evaluator uses the real catalog. |
| 6 | No API section | Zero token usage, no hosted model, and teammates can run it with ps1 or sh scripts. |
14. Remaining Tasks
| Task | Owner | Status |
|---|---|---|
Confirm no open final PR exists before final packaging from main. | Team / @yinasaurus | No open PRs at last check |
| Confirm individual member names and contribution split. | Team | Filled in README and Devpost draft |
| Record and upload public YouTube video. | Demo owner | Still needed |
| Paste Devpost draft and links. | Submitter | Still needed |
Build final zip plus .zip.sha256 checksum. | Engineering | Run package script from latest main |
| Keep score above 0.80 and avoid paid APIs. | Engineering | Currently satisfied at 0.955300 |