Our Data Pipeline
RedRecs operates through a systematic data pipeline that analyzes Reddit discussions to surface authentic product reviews from real users—filtering out influencer content and press releases to deliver genuine community insights.
The Process:
Identify product categories and relevant discussions
Extract and analyze user reviews
Map feedback to specific product models
Calculate sentiment-based rankings
Verify and refine data accuracy
Pipeline Overview
1. Thread Collection
RedRecs aggregates relevant Reddit discussions from the past year across target product categories.
The system generates category-specific queries (e.g., "best air purifier recommendations") and evaluates each thread for relevance using AI validation. Collection continues for each query while relevance scores remain above 40%, ensuring efficient data gathering focused on quality discussions.
2. Review Extraction
Once threads are collected, the system identifies users who have posted substantive product reviews—not just casual mentions.
For each reviewer, RedRecs captures:
Complete discussion context (subreddit, post, and replies)
Reddit username and overall sentiment
Product identification (brand, model, specifications)
Product URLs where available
Key quotes highlighting specific feedback
Large threads are intelligently segmented to maintain accuracy while preserving comment thread integrity.
3. Product Model Mapping
This stage addresses a core challenge: Reddit users rarely use precise model names. Comments like "GPX 2 mouse" or "Ninja 6-in-1 dual basket" require interpretation to identify exact products.
Our Approach: A web research agent cross-references each product mention with Google search results to identify all potential matching models. Results are cached to prevent redundant research.
Model Identification: RedRecs uses model name combined with descriptive specifications as unique identifiers. When a new model is discovered, the system checks for existing matches through:
String matching algorithms
AI-powered comparison of names and specifications
Once identified, all associated retail links (Amazon, Walmart, brand sites) are automatically connected to that product model.
4. Sentiment-Based Ranking
Rankings reflect aggregated community sentiment, highlighting the most recommended products on Reddit.
Ranking Factors:
Volume of positive user mentions
Positive-to-negative feedback ratio
Confidence level of product identification
Scoring Methodology:
Each user contributes one vote per product (regardless of comment frequency)
Ambiguous mentions are distributed proportionally across potential matches
Popular models receive measured weighting adjustments
The final Community Score combines normalized positive sentiment (75%) with the positive-to-negative ratio (25%).
5. Quality Assurance
Data undergoes manual review through an internal dashboard, enabling efficient verification and cleanup without direct database manipulation.
The system also groups product variants into series when appropriate. For example, when users commonly reference "Ninja grill" without specifying a model, related variants are consolidated into a "Ninja Grill Series" to prevent ranking fragmentation.
Technology Stack
AI & Language Models
OpenAI (GPT-4o, o3-mini)
Google Gemini (2.5 Flash)
Data Sources
Reddit API (PRAW)
Google Search API
Amazon Product Advertising API
BrightData (e-commerce scraping)
FireCrawl + Jina AI (web scraping)
Perplexity (supplementary research)
Infrastructure
Python (automation scripts)
HTML/JavaScript/TypeScript/Nuxt (frontend)
Supabase (database)
Cursor (development environment)
Replit (script deployment)
Cloudflare Pages (hosting)