Browse Categories
All

How RedRecs works

Last updated: October 27, 2025

Our Data Pipeline

RedRecs operates through a systematic data pipeline that analyzes Reddit discussions to surface authentic product reviews from real users—filtering out influencer content and press releases to deliver genuine community insights.

The Process:

  • Identify product categories and relevant discussions

  • Extract and analyze user reviews

  • Map feedback to specific product models

  • Calculate sentiment-based rankings

  • Verify and refine data accuracy

Pipeline Overview

1. Thread Collection

RedRecs aggregates relevant Reddit discussions from the past year across target product categories.

The system generates category-specific queries (e.g., "best air purifier recommendations") and evaluates each thread for relevance using AI validation. Collection continues for each query while relevance scores remain above 40%, ensuring efficient data gathering focused on quality discussions.

2. Review Extraction

Once threads are collected, the system identifies users who have posted substantive product reviews—not just casual mentions.

For each reviewer, RedRecs captures:

  • Complete discussion context (subreddit, post, and replies)

  • Reddit username and overall sentiment

  • Product identification (brand, model, specifications)

  • Product URLs where available

  • Key quotes highlighting specific feedback

Large threads are intelligently segmented to maintain accuracy while preserving comment thread integrity.

3. Product Model Mapping

This stage addresses a core challenge: Reddit users rarely use precise model names. Comments like "GPX 2 mouse" or "Ninja 6-in-1 dual basket" require interpretation to identify exact products.

Our Approach: A web research agent cross-references each product mention with Google search results to identify all potential matching models. Results are cached to prevent redundant research.

Model Identification: RedRecs uses model name combined with descriptive specifications as unique identifiers. When a new model is discovered, the system checks for existing matches through:

  • String matching algorithms

  • AI-powered comparison of names and specifications

Once identified, all associated retail links (Amazon, Walmart, brand sites) are automatically connected to that product model.

4. Sentiment-Based Ranking

Rankings reflect aggregated community sentiment, highlighting the most recommended products on Reddit.

Ranking Factors:

  • Volume of positive user mentions

  • Positive-to-negative feedback ratio

  • Confidence level of product identification

Scoring Methodology:

  • Each user contributes one vote per product (regardless of comment frequency)

  • Ambiguous mentions are distributed proportionally across potential matches

  • Popular models receive measured weighting adjustments

The final Community Score combines normalized positive sentiment (75%) with the positive-to-negative ratio (25%).

5. Quality Assurance

Data undergoes manual review through an internal dashboard, enabling efficient verification and cleanup without direct database manipulation.

The system also groups product variants into series when appropriate. For example, when users commonly reference "Ninja grill" without specifying a model, related variants are consolidated into a "Ninja Grill Series" to prevent ranking fragmentation.

Technology Stack

AI & Language Models

  • OpenAI (GPT-4o, o3-mini)

  • Google Gemini (2.5 Flash)

Data Sources

  • Reddit API (PRAW)

  • Google Search API

  • Amazon Product Advertising API

  • BrightData (e-commerce scraping)

  • FireCrawl + Jina AI (web scraping)

  • Perplexity (supplementary research)

Infrastructure

  • Python (automation scripts)

  • HTML/JavaScript/TypeScript/Nuxt (frontend)

  • Supabase (database)

  • Cursor (development environment)

  • Replit (script deployment)

  • Cloudflare Pages (hosting)