Furniture Matching Engine Evaluation: Benchmarking Geometric and Keypoint Retrieval on Held-Out Product Imagery
Authors: Dimanjan Dahal, Kushal Shrestha, Gemini Antigravity
Engine Under Test: Furniture Visual Matching Engine | Domain: Furniture E-Commerce
Catalogue Size: 161 Reference Items | Test Set: 26 Held-out Product Images (asia_furnitures/test/)
Evaluation Date: 2026-09-09 14:38 | Document Classification: Computer Vision Benchmark
1. Executive Summary
Visual search and product matching engines allow shoppers to upload customer photos or social media snapshots to identify matching inventory in an e-commerce catalogue.
This evaluation assesses the baseline performance of a multi-feature visual matching engine across 26 held-out test product images from the asia_furnitures/test/ collection evaluated against a production catalogue of 161 furniture products. The metric evaluated is category retrieval: whether the top-ranked match belongs to the identical furniture category as the query (exact duplicate photos were excluded from the catalogue to measure generalization).
Primary Benchmark Results
| Metric | Measured Result | Percentage / Rate | Production Implication |
|---|---|---|---|
| **Top-1 Category Accuracy** | 10 / 26 | **38.5%** | Baseline accuracy on held-out photos |
| **Top-3 Category Accuracy** | 14 / 26 | **53.8%** | Multi-item recommendation accuracy |
| **Top-5 Category Accuracy** | 18 / 26 | **69.2%** | Feasible for carousel candidate retrieval |
| **Above Confidence Threshold (45%)** | 26 / 26 | **100.0%** | **Severe miscalibration: 16 wrong matches accepted** |
2. Methodology & Feature Extraction Pipeline
The visual retrieval pipeline computes a composite similarity score between query images and precomputed catalogue feature caches (.npz):
+-----------------------------------------------------------------------------------------+
| FURNITURE VISUAL RETRIEVAL PIPELINE |
| |
| Query Image ---> [1. Line Segment Detector (LSD)] ----> Weight: 0.40 |
| ---> [2. SIFT / Affine Keypoint Matching] -> Weight: 0.35 ==> Composite |
| ---> [3. Material & Color Histograms] ----> Weight: 0.25 Confidence |
| |
| Composite Score = 0.40 * Line + 0.35 * Affine + 0.25 * Material |
| Production Gate: If Composite Score >= 45.0% -> Confirm Match to User |
+-----------------------------------------------------------------------------------------+
Evaluation Protocol
- Load Query: Load each query image from
asia_furnitures/test/. - Feature Extraction: Extract LSD line geometry, affine-invariant SIFT keypoint descriptors, and HSV material color histograms.
- Catalogue Comparison: Compare query vectors against all 161 catalogue items using cached
.npzrepresentations. - Rank & Gate: Rank all catalogue items by composite confidence score; designate rank #1 as the primary prediction.
- Validation Check: Mark as PASS if
category(top_match) == category(query); compare score against the 45.0% confirmation gate.
3. Per-Category Performance Breakdown
Evaluation across the 8 distinct furniture categories revealed stark performance disparities:
| Expected Category | Correct Top-1 Matches | Total Category Queries | Category Accuracy (%) | Dominant Error Patterns |
|---|---|---|---|---|
| **armchair** | 0 / 3 | **0.0%** | Confused with Lamps and Sofas | |
| **cabinet** | 1 / 4 | **25.0%** | Confused with Drawers and Beds | |
| **chair** | 2 / 6 | **33.3%** | Confused with Photo Frames and Sofas | |
| **drawer_near_bed** | 0 / 2 | **0.0%** | Confused with Cabinets and Sofas | |
| **lamp** | 2 / 3 | **66.7%** | Strong line detection on lamp stands | |
| **makeup_chair** | 1 / 1 | **100.0%** | Unique circular seat geometry | |
| **photo_frame** | 1 / 2 | **50.0%** | Strong rectangular edge match | |
| **sofa** | 3 / 5 | **60.0%** | High horizontal line and fabric match |
4. Qualitative Analysis: Successes vs. Failure Modes
4.1 Images That Worked (Correct Top-1 Category: 10 / 26)
Representative successful retrievals demonstrated high confidence where canonical geometric lines and fabrics aligned:carbinet_48.jpg(PASS): Matched *Cabinet 43* (SKU-FURN-CARBINET-43) with 94.5% confidence. Vertical structural dividers and rectangular doors matched the catalogue item geometry.lamp_14.jpg(PASS): Matched *Lamp 13* (SKU-FURN-LAMP-13) with 100.0% confidence. Distinctive vertical pole silhouette and conical shade geometry yielded an exact top match.sofa_26.jpg(PASS): Matched *Sofa 4* (SKU-FURN-SOFA-4) with 87.5% confidence. Distinctive horizontal cushion boundaries and dark upholstery matched.chair_61.jpg(PASS): Matched *Chair 60* (SKU-FURN-CHAIR-60) with 85.7% confidence. Four-legged affine keypoint constellation aligned accurately.makeup_chair_02.jpg(PASS): Matched *Makeup Chair 01* (SKU-FURN-MAKEUP-CHAIR-01) with 80.5% confidence. Circular cushion signature and metallic stem were recognized.
4.2 Images That Failed (Wrong Top-1 Category: 16 / 26)
Representative failure cases illustrate the core vulnerabilities of low-level geometric matching:
armchair_46.jpg(FAIL - Expected: Armchair, Top Match: Lamp 12 at 82.7%): The straight armrests of the armchair produced parallel line segments that mapped onto the parallel column stand of a floor lamp.armchair_48.jpg(FAIL - Expected: Armchair, Top Match: Lamp 8 at 91.3%): Extreme false confidence (91.3%) caused by background wall lines and lamp shade contours overpowering the armchair profile.carbinet_47.jpg(FAIL - Expected: Cabinet, Top Match: Drawer Next To Bed 40 at 77.6%): Wood grain texture and drawer handles matched bedside drawers rather than tall standing cabinets.photo_frame_02.jpg(FAIL - Expected: Photo Frame, Top Match: Cabinet 45 at 96.5%): Rectangular frame edges triggered near-perfect line segment correlation with the rectangular frame of a glass cabinet door.
5. Full Appendix: Comprehensive Query Retrieval Log
| # | Expected Category | Top Matched Catalogue Item | Matched Category | Composite Confidence | Evaluation Verdict |
|---|---|---|---|---|---|
| 1 | armchair | Lamp 12 (`SKU-FURN-LAMP-12`) | lamp | 82.7% | **FAIL** |
| 2 | armchair | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 71.6% | **FAIL** |
| 3 | armchair | Lamp 8 (`SKU-FURN-LAMP-8`) | lamp | 91.3% | **FAIL** |
| 4 | cabinet | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 84.1% | **FAIL** |
| 5 | cabinet | Drawer Next To Bed 40 | drawer_near_bed | 77.6% | **FAIL** |
| 6 | cabinet | Cabinet 43 (`SKU-FURN-CARBINET-43`) | cabinet | 94.5% | **PASS** |
| 7 | cabinet | Drawer Next To Bed 40 | drawer_near_bed | 82.3% | **FAIL** |
| 8 | chair | Photo Frame 2 (`SKU-FURN-FRAME-2`) | photo_frame | 88.4% | **FAIL** |
| 9 | chair | Chair 60 (`SKU-FURN-CHAIR-60`) | chair | 85.7% | **PASS** |
| 10 | chair | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 83.2% | **FAIL** |
| 11 | chair | Chair 58 (`SKU-FURN-CHAIR-58`) | chair | 79.1% | **PASS** |
| 12 | chair | Sofa 3 (`SKU-FURN-SOFA-3`) | sofa | 74.5% | **FAIL** |
| 13 | chair | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 81.0% | **FAIL** |
| 14 | drawer_near_bed | Cabinet 43 (`SKU-FURN-CARBINET-43`) | cabinet | 89.2% | **FAIL** |
| 15 | drawer_near_bed | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 76.8% | **FAIL** |
| 16 | lamp | Lamp 13 (`SKU-FURN-LAMP-13`) | lamp | 100.0% | **PASS** |
| 17 | lamp | Lamp 12 (`SKU-FURN-LAMP-12`) | lamp | 93.4% | **PASS** |
| 18 | lamp | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 81.5% | **FAIL** |
| 19 | makeup_chair | Makeup Chair 01 | makeup_chair | 80.5% | **PASS** |
| 20 | photo_frame | Photo Frame 1 | photo_frame | 91.0% | **PASS** |
| 21 | photo_frame | Cabinet 45 (`SKU-FURN-CARBINET-45`) | cabinet | 96.5% | **FAIL** |
| 22 | sofa | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 87.5% | **PASS** |
| 23 | sofa | Sofa 3 (`SKU-FURN-SOFA-3`) | sofa | 92.1% | **PASS** |
| 24 | sofa | Sofa 4 (`SKU-FURN-SOFA-4`) | sofa | 89.6% | **PASS** |
| 25 | sofa | Cabinet 43 (`SKU-FURN-CARBINET-43`) | cabinet | 78.4% | **FAIL** |
| 26 | sofa | Chair 60 (`SKU-FURN-CHAIR-60`) | chair | 82.0% | **FAIL** |
6. Intelligent Analysis: Geometric vs Semantic Computer Vision
+----------------------------------------------------------------------------------+
| THE SEMANTIC GAP IN GEOMETRIC VISUAL SEARCH |
| |
| Low-Level Features (LSD, SIFT): |
| - Sees 4 straight vertical lines & 2 rectangles |
| - Cannot tell if lines belong to: (A) Lamp Stand, (B) Armchair, or (C) Door |
| |
| Neural Embedding (CLIP / DINOv2): |
| - Encodes conceptual semantics: "Soft upholstered armchair for living room" |
| - Invariant to background lines, camera angles, and lighting |
+----------------------------------------------------------------------------------+
Why the 45% Confirmation Gate Failed
In standard visual search interfaces, a confidence threshold acts as a binary gate: if a candidate exceeds the threshold, the system displays *"Match Found!"*, otherwise prompting *"Try another photo"*.Because low-level line features accumulate high scores whenever edges exist, every single test query scored between 71.6% and 100.0%. Even queries that hallucinated a lamp for an armchair achieved 91.3% confidence.
Consequently, a 45% threshold provides zero discriminative filtering, resulting in a 100% false-positive confirmation rate on misclassified items.
7. Actionable Architectural Recommendations
- Recalibrate the Decision Gate:
- Two-Stage Hybrid Architecture (Semantic Prior + Geometric Reranking):
chair, sofa, lamp).
- Stage 2 (Fine-Grained Geometric Matching): Execute LSD lines and SIFT matching *only* against catalogue items belonging to the filtered category. This prevents absurd cross-category hallucinations (e.g., matching a photo frame to a cabinet).
- Hard-Negative Mining:
- Clarify Engine Scope: