CLEARCART  ·  Source Attribution Pipeline  ·  Internal Tool

Data Pipeline — Evidence Ingestion

Turning creator reviews into credited evidence

CLEARCART's five-pillar scores need evidence, and hands-on testing every product in every category isn't feasible for a pilot. This pipeline takes content-creator review sites — the people already buying, testing, and disassembling products — profiles how trustworthy their methodology is, and extracts their claims into structured, pillar-mapped, confidence-rated records with mandatory attribution back to the original reviewer.

01

Fetch a source's homepage, methodology page, and product review pages (or YouTube transcript)

02

Profile the reviewer's methodology, monetization, and pillar coverage into a trust tier

03

Extract per-product claims, typed by evidence kind, mapped to the five pillars

04

Flag every pillar the reviewer didn't cover as an explicit gap, not a silent zero

05

Emit attributed JSON records ready for the dashboard

Step 01 extraction — standard web pages, JS-rendered tables (e.g. Clarity Insight's swaps database), and YouTube video transcripts alike — is powered by Tavily.

What this pipeline refuses to do

Claims, not scores, are ingested

A reviewer's overall verdict (e.g. "9/10") is stored as context only. What actually feeds the pillars is individual claims, each typed by evidence kind — CLEARCART does its own scoring on top of that.

Gaps are first-class data

Every pillar a reviewer skips is recorded as an explicit gap with a note on what CLEARCART must source independently — not left blank or implicitly scored as zero. That gap record is itself the differentiation story.

Attribution is mandatory

Every record carries the reviewer's name and source URL so the dashboard — and eventually the extension — can credit exactly whose testing backs a claim.

Confidence caps by evidence kind

Anecdotes cap at low, third-party citations cap at medium until independently verified, and only instrumented testing reaches high — regardless of how confidently the reviewer states it.

Three reviewers, three trust tiers

The same profiling step applied to three very different kinds of review sites — showing why trust tier can't be assumed from a source's popularity or polish alone.

Prudent Reviews

High Trust hands_on_testing

prudentreviews.com  ·  monetization: free samples (declined by default) + newsletter

Physically buys and uses each product for weeks to years — real-world use tests plus instrumented experiments (heat-conduction and heat-retention timing with surface thermometers). Supplements with expert interviews for categories impractical to fully test, and retains products as long-term benchmarks against new competitors.

Conflict of interest: discloses that most free-sample offers from brands are declined, and reviews accepted samples the same as purchased ones, including reporting negative findings. No explicit affiliate disclosure surfaced in the fetched text — footers are stripped before analysis, so this should be verified directly on the site.

PillarCoverageNote
🌱 PlanetAbsentNo environmental impact or materials-footprint content found.
🤝 PeoplePartialReports country of manufacture and raw-material regions, but no labor conditions or factory audits.
⭐ QualityStrongInstrumented, repeatable testing across 15+ competing products.
💰 ValueStrongExplicit price-vs-competitor positioning tied directly to test results.
🔍 TransparencyStrongDedicated methodology page, named reviewer, published raw data tables.

Clarity Insight Advisory (Buy Better With Brian)

Medium Trust desk_research

clarityinsightadvisory.com  ·  monetization: affiliate links + newsletter

A former Wall Street credit trader applying a risk-management framework: checks ingredients and certifications against primary regulatory/toxicology sources rather than hands-on testing. Weighs cumulative exposure, no-safe-dose chemical categories, the gap between "approved" and "proven safe," and downstream environmental impact — scoring 0–10 against cheaper-or-equal, lower-risk alternatives.

Conflict of interest: discloses that affiliate links are attached only after a product is scored ("the research picks the product, then the link gets attached") and states $0 in brand sponsorships — a self-reported claim of independence, not an independently audited one.

PillarCoverageNote
🌱 PlanetPartialCovers downstream water contamination (e.g. PFAS in wastewater) but not manufacturing emissions, packaging, or biodegradability.
🤝 PeopleAbsentNo discussion of labor conditions, factory practices, or sourcing ethics.
⭐ QualityAbsentOriented around toxicological risk, not performance or durability.
💰 ValuePartialCompares price against risk-avoidance only, not broader factors like longevity.
🔍 TransparencyStrongExplicit 0–10 rubric, named individual, disclosed monetization model.

Fineas Jackson

Low Trust desk_research (video)

youtube.com/@FineasJackson  ·  monetization: none disclosed

Explains the engineering and material differences behind quality signals (zipper tooth formation, tape density, slider alloy; down fill-power and moisture handling) using historical and technical context — without a disclosed protocol for how products are selected, acquired, or personally verified.

Conflict of interest: no monetization, sponsorship, or affiliate disclosure found in the fetched content, so conflicts can't be assessed from what's available. A "Brand: fineasjackson.com" link on the channel page wasn't captured in the fetched text and is worth checking separately.

PillarCoverageNote
🌱 PlanetAbsentNo environmental or sustainability content in sampled videos.
🤝 PeoplePartialNames countries of origin for premium hardware as a quality signal; doesn't address labor conditions directly.
⭐ QualityStrongDetailed technical breakdown of cheap-vs-premium construction and materials.
💰 ValueAbsentNo price-to-quality comparisons in sampled content.
🔍 TransparencyAbsentNo disclosed methodology or rubric — educational rather than a disclosed rating system.

Six products, pillar by pillar

Per-product status shows whether each pillar has directly evidenced claims, brand-claim-only support, or is an open gap this source didn't cover.

Product Source 🌱 Planet 🤝 People ⭐ Quality 💰 Value 🔍 Transparency
Made In Stainless Clad Frying PanMade In · cookware Prudent Reviews ○ Gap ◐ Brand claim ● Evidenced ● Evidenced ● Evidenced
de Buyer Mineral B Pro Carbon Steelde Buyer · cookware Clarity Insight ◐ Brand claim ○ Gap ● Evidenced ● Evidenced ◐ Brand claim
Naturepedic EOS Classic MattressNaturepedic · mattresses Clarity Insight ◐ Brand claim ◐ Brand claim ● Evidenced ● Evidenced ● Evidenced
AspenClean Laundry Powder + BoosterAspenClean · laundry detergent Clarity Insight ◐ Brand claim ○ Gap ○ Gap ● Evidenced ● Evidenced
Schlage Encode Plus Smart LockSchlage · smart locks Clarity Insight ○ Gap ○ Gap ● Evidenced ● Evidenced ◐ Brand claim
Feathered Friends Bavarian 850 Fill ComforterFeathered Friends · bedding Fineas Jackson ○ Gap ○ Gap ● Evidenced ○ Gap ○ Gap
Evidenced — tested or independently cited claim Brand claim only — repeated, not independently verified Gap — this source didn't cover the pillar; CLEARCART must source independently

Evidence kind sets a confidence ceiling

Whatever confidence a reviewer implies, this pipeline caps it by how the claim was actually substantiated — matching the same three-tier confidence system (High/Medium/Low) used by the extension's scoring engine (pillars.js), so pipeline output plugs directly into pillar scoring without re-mapping.

Evidence kindConfidence capExample
testedHighPrudent Reviews' instrumented boil-time and heat-retention measurements
cited_third_partyMedium (until verified)Clarity Insight citing a University of Newcastle 2022 microplastics study
brand_claim_repeatedLow"Handcrafted by Amish artisans in an Ohio factory" — Naturepedic's own claim, repeated by the reviewer
anecdotalLowOwner reports from r/carbonsteel on long-term seasoning behavior