Computer Vision & Retrieval

Furniture Matching Engine Evaluation: Benchmarking Geometric and Keypoint Retrieval on Held-Out Product Imagery

An empirical evaluation of a multi-signal visual retrieval engine (LSD line geometry, SIFT/affine keypoints, material signature) on 26 held-out furniture images against a 161-item catalogue. While top-5 retrieval reaches 69.2%, top-1 category accuracy is 38.5% and a 45% confidence threshold creates a 100% false-pass rate.

Dimanjan Dahal, Kushal Shrestha, Gemini Antigravity
13 min read
9/9/2026

Furniture Matching Engine Evaluation: Benchmarking Geometric and Keypoint Retrieval on Held-Out Product Imagery

Authors: Dimanjan Dahal, Kushal Shrestha, Gemini Antigravity
Engine Under Test: Furniture Visual Matching Engine | Domain: Furniture E-Commerce
Catalogue Size: 161 Reference Items | Test Set: 26 Held-out Product Images (asia_furnitures/test/)
Evaluation Date: 2026-09-09 14:38 | Document Classification: Computer Vision Benchmark


Key Architectural Takeaway: Evaluating low-level geometric and keypoint matching (LSD lines + SIFT/affine keypoints + color-material histograms) on non-duplicate product photos revealed that top-1 category accuracy is only 38.5% (10/26), though top-5 candidate recall reaches 69.2%. Crucially, the production confidence threshold of 45.0% was exceeded by 100% of queries (including all 16 incorrect matches), exposing a critical confidence calibration defect. Pure geometric pipelines fail on open-set visual search without a semantic neural prior (e.g., CLIP / DINOv2).

1. Executive Summary

Visual search and product matching engines allow shoppers to upload customer photos or social media snapshots to identify matching inventory in an e-commerce catalogue.

This evaluation assesses the baseline performance of a multi-feature visual matching engine across 26 held-out test product images from the asia_furnitures/test/ collection evaluated against a production catalogue of 161 furniture products. The metric evaluated is category retrieval: whether the top-ranked match belongs to the identical furniture category as the query (exact duplicate photos were excluded from the catalogue to measure generalization).

Primary Benchmark Results

MetricMeasured ResultPercentage / RateProduction Implication
**Top-1 Category Accuracy**10 / 26**38.5%**Baseline accuracy on held-out photos
**Top-3 Category Accuracy**14 / 26**53.8%**Multi-item recommendation accuracy
**Top-5 Category Accuracy**18 / 26**69.2%**Feasible for carousel candidate retrieval
**Above Confidence Threshold (45%)**26 / 26**100.0%****Severe miscalibration: 16 wrong matches accepted**

2. Methodology & Feature Extraction Pipeline

The visual retrieval pipeline computes a composite similarity score between query images and precomputed catalogue feature caches (.npz):

+-----------------------------------------------------------------------------------------+
|                        FURNITURE VISUAL RETRIEVAL PIPELINE                              |
|                                                                                         |
|  Query Image ---> [1. Line Segment Detector (LSD)] ----> Weight: 0.40                   |
|              ---> [2. SIFT / Affine Keypoint Matching] -> Weight: 0.35  ==> Composite   |
|              ---> [3. Material & Color Histograms] ----> Weight: 0.25      Confidence    |
|                                                                                         |
|  Composite Score = 0.40 * Line + 0.35 * Affine + 0.25 * Material                        |
|  Production Gate: If Composite Score >= 45.0% -> Confirm Match to User                   |
+-----------------------------------------------------------------------------------------+

Evaluation Protocol

  1. Load Query: Load each query image from asia_furnitures/test/.
  2. Feature Extraction: Extract LSD line geometry, affine-invariant SIFT keypoint descriptors, and HSV material color histograms.
  3. Catalogue Comparison: Compare query vectors against all 161 catalogue items using cached .npz representations.
  4. Rank & Gate: Rank all catalogue items by composite confidence score; designate rank #1 as the primary prediction.
  5. Validation Check: Mark as PASS if category(top_match) == category(query); compare score against the 45.0% confirmation gate.

3. Per-Category Performance Breakdown

Evaluation across the 8 distinct furniture categories revealed stark performance disparities:

Expected CategoryCorrect Top-1 MatchesTotal Category QueriesCategory Accuracy (%)Dominant Error Patterns
**armchair**0 / 3**0.0%**Confused with Lamps and Sofas
**cabinet**1 / 4**25.0%**Confused with Drawers and Beds
**chair**2 / 6**33.3%**Confused with Photo Frames and Sofas
**drawer_near_bed**0 / 2**0.0%**Confused with Cabinets and Sofas
**lamp**2 / 3**66.7%**Strong line detection on lamp stands
**makeup_chair**1 / 1**100.0%**Unique circular seat geometry
**photo_frame**1 / 2**50.0%**Strong rectangular edge match
**sofa**3 / 5**60.0%**High horizontal line and fabric match

4. Qualitative Analysis: Successes vs. Failure Modes

4.1 Images That Worked (Correct Top-1 Category: 10 / 26)

Representative successful retrievals demonstrated high confidence where canonical geometric lines and fabrics aligned:
  • carbinet_48.jpg (PASS): Matched *Cabinet 43* (SKU-FURN-CARBINET-43) with 94.5% confidence. Vertical structural dividers and rectangular doors matched the catalogue item geometry.
  • lamp_14.jpg (PASS): Matched *Lamp 13* (SKU-FURN-LAMP-13) with 100.0% confidence. Distinctive vertical pole silhouette and conical shade geometry yielded an exact top match.
  • sofa_26.jpg (PASS): Matched *Sofa 4* (SKU-FURN-SOFA-4) with 87.5% confidence. Distinctive horizontal cushion boundaries and dark upholstery matched.
  • chair_61.jpg (PASS): Matched *Chair 60* (SKU-FURN-CHAIR-60) with 85.7% confidence. Four-legged affine keypoint constellation aligned accurately.
  • makeup_chair_02.jpg (PASS): Matched *Makeup Chair 01* (SKU-FURN-MAKEUP-CHAIR-01) with 80.5% confidence. Circular cushion signature and metallic stem were recognized.

4.2 Images That Failed (Wrong Top-1 Category: 16 / 26)


Representative failure cases illustrate the core vulnerabilities of low-level geometric matching:
  • armchair_46.jpg (FAIL - Expected: Armchair, Top Match: Lamp 12 at 82.7%): The straight armrests of the armchair produced parallel line segments that mapped onto the parallel column stand of a floor lamp.

  • armchair_48.jpg (FAIL - Expected: Armchair, Top Match: Lamp 8 at 91.3%): Extreme false confidence (91.3%) caused by background wall lines and lamp shade contours overpowering the armchair profile.

  • carbinet_47.jpg (FAIL - Expected: Cabinet, Top Match: Drawer Next To Bed 40 at 77.6%): Wood grain texture and drawer handles matched bedside drawers rather than tall standing cabinets.

  • photo_frame_02.jpg (FAIL - Expected: Photo Frame, Top Match: Cabinet 45 at 96.5%): Rectangular frame edges triggered near-perfect line segment correlation with the rectangular frame of a glass cabinet door.

5. Full Appendix: Comprehensive Query Retrieval Log

#Expected CategoryTop Matched Catalogue ItemMatched CategoryComposite ConfidenceEvaluation Verdict
1armchairLamp 12 (`SKU-FURN-LAMP-12`)lamp82.7%**FAIL**
2armchairSofa 4 (`SKU-FURN-SOFA-4`)sofa71.6%**FAIL**
3armchairLamp 8 (`SKU-FURN-LAMP-8`)lamp91.3%**FAIL**
4cabinetSofa 4 (`SKU-FURN-SOFA-4`)sofa84.1%**FAIL**
5cabinetDrawer Next To Bed 40drawer_near_bed77.6%**FAIL**
6cabinetCabinet 43 (`SKU-FURN-CARBINET-43`)cabinet94.5%**PASS**
7cabinetDrawer Next To Bed 40drawer_near_bed82.3%**FAIL**
8chairPhoto Frame 2 (`SKU-FURN-FRAME-2`)photo_frame88.4%**FAIL**
9chairChair 60 (`SKU-FURN-CHAIR-60`)chair85.7%**PASS**
10chairSofa 4 (`SKU-FURN-SOFA-4`)sofa83.2%**FAIL**
11chairChair 58 (`SKU-FURN-CHAIR-58`)chair79.1%**PASS**
12chairSofa 3 (`SKU-FURN-SOFA-3`)sofa74.5%**FAIL**
13chairSofa 4 (`SKU-FURN-SOFA-4`)sofa81.0%**FAIL**
14drawer_near_bedCabinet 43 (`SKU-FURN-CARBINET-43`)cabinet89.2%**FAIL**
15drawer_near_bedSofa 4 (`SKU-FURN-SOFA-4`)sofa76.8%**FAIL**
16lampLamp 13 (`SKU-FURN-LAMP-13`)lamp100.0%**PASS**
17lampLamp 12 (`SKU-FURN-LAMP-12`)lamp93.4%**PASS**
18lampSofa 4 (`SKU-FURN-SOFA-4`)sofa81.5%**FAIL**
19makeup_chairMakeup Chair 01makeup_chair80.5%**PASS**
20photo_framePhoto Frame 1photo_frame91.0%**PASS**
21photo_frameCabinet 45 (`SKU-FURN-CARBINET-45`)cabinet96.5%**FAIL**
22sofaSofa 4 (`SKU-FURN-SOFA-4`)sofa87.5%**PASS**
23sofaSofa 3 (`SKU-FURN-SOFA-3`)sofa92.1%**PASS**
24sofaSofa 4 (`SKU-FURN-SOFA-4`)sofa89.6%**PASS**
25sofaCabinet 43 (`SKU-FURN-CARBINET-43`)cabinet78.4%**FAIL**
26sofaChair 60 (`SKU-FURN-CHAIR-60`)chair82.0%**FAIL**

6. Intelligent Analysis: Geometric vs Semantic Computer Vision

+----------------------------------------------------------------------------------+
|                   THE SEMANTIC GAP IN GEOMETRIC VISUAL SEARCH                    |
|                                                                                  |
|  Low-Level Features (LSD, SIFT):                                                |
|  - Sees 4 straight vertical lines & 2 rectangles                                 |
|  - Cannot tell if lines belong to: (A) Lamp Stand, (B) Armchair, or (C) Door     |
|                                                                                  |
|  Neural Embedding (CLIP / DINOv2):                                               |
|  - Encodes conceptual semantics: "Soft upholstered armchair for living room"     |
|  - Invariant to background lines, camera angles, and lighting                     |
+----------------------------------------------------------------------------------+

Why the 45% Confirmation Gate Failed

In standard visual search interfaces, a confidence threshold acts as a binary gate: if a candidate exceeds the threshold, the system displays *"Match Found!"*, otherwise prompting *"Try another photo"*.

Because low-level line features accumulate high scores whenever edges exist, every single test query scored between 71.6% and 100.0%. Even queries that hallucinated a lamp for an armchair achieved 91.3% confidence.

Consequently, a 45% threshold provides zero discriminative filtering, resulting in a 100% false-positive confirmation rate on misclassified items.


7. Actionable Architectural Recommendations

  1. Recalibrate the Decision Gate:
- Immediately raise the confidence threshold from 45.0% to 85.0%. - Implement relative margin gating: require that the #1 match exceeds the #2 match by at least 12.0% confidence margin before confirming an automated match.
  1. Two-Stage Hybrid Architecture (Semantic Prior + Geometric Reranking):
- Stage 1 (Coarse Semantic Filter): Run query images through a lightweight zero-shot classifier (CLIP ViT-B/32 or MobileNetV4) to determine the top-1 coarse category (chair, sofa, lamp). - Stage 2 (Fine-Grained Geometric Matching): Execute LSD lines and SIFT matching *only* against catalogue items belonging to the filtered category. This prevents absurd cross-category hallucinations (e.g., matching a photo frame to a cabinet).
  1. Hard-Negative Mining:
- Noticeably, *Sofa 4* and *Cabinet 43* appeared repeatedly as false top matches across multiple categories. Downweight frequently matching hub items by applying Inverse Document Frequency (IDF) penalties across keypoint clusters.
  1. Clarify Engine Scope:
- The geometric engine is optimized for near-duplicate re-identification (identifying the exact same physical product under different lighting or angles), not open-set semantic categorization. Product marketing and UI copy should align with this technical boundary.
Tags:
furniture matchingvisual searchcomputer visiongeometric matchingLSD line geometrySIFT keypointsaffine matchingmaterial signaturee-commerce retrievalconfidence calibration

Original Document & Benchmark Telemetry

Download the full raw experimental investigation report in PDF format.

Download PDF Report

Ready to Implement AI Solutions?

Based on this research, let Sajedar help you build conversational AI solutions tailored for the Nepal and South Asia market.

Chat with us