AI/ML

Fighting Misinformation with Machine Learning

Building a fake news detection pipeline using NLP — text vectorization, model selection, and the challenges of training on real-world data.

May 5, 20269 min read
Suyash Vakhariya
Suyash VakhariyaAI Engineer & Technical Product Manager

Why This Matters

Misinformation spreads 6x faster than factual news on social media. Automated detection systems aren't a silver bullet, but they're an essential tool in the fight against fake news. Here's how I built one.

The Dataset

I used a dataset of ~44,000 news articles, labeled as "real" or "fake". Each article has:

  • Title: headline text
  • Body: full article text
  • Label: real (0) or fake (1)

The key insight: combining title and body text significantly improves classification accuracy, because fake news often has sensational headlines that don't match the article's substance.

python
df['text'] = (df['title'].fillna('') + ' ' + df['text'].fillna('')).str.strip()

The Pipeline

1. Text Preprocessing

python
from sklearn.feature_extraction.text import TfidfVectorizer

tfidf = TfidfVectorizer(
    lowercase=True,
    stop_words='english',
    max_df=0.95,   # Remove words in >95% of articles (too common)
    min_df=5        # Remove words in <5 articles (too rare)
)

2. Model Comparison

ModelAccuracyNotes
Logistic Regression93.8%Best balance of speed and accuracy
Random Forest91.2%Slower, prone to overfitting
Naive Bayes89.5%Very fast, but lower accuracy
Passive Aggressive94.1%Best accuracy, less interpretable

I chose Logistic Regression for the final model because:

  • Near-best accuracy (93.8%)
  • Highly interpretable (you can see which words drive predictions)
  • Fast inference (important for real-time classification)
  • Low memory footprint

3. What the Model Learned

The most informative features reveal clear patterns:

Strong fake indicators: "breaking", "shocking", "you won't believe", "share this", "urgent", "exposed"

Strong real indicators: "according to", "officials said", "reported", "study", "data shows", "percent"

This makes intuitive sense — fake news uses emotional, sensational language, while real reporting uses attribution and evidence-based language.

The Hard Part: Distribution Shift

The biggest challenge isn't building the model — it's keeping it accurate over time. News language evolves. New topics emerge. Political vocabulary shifts.

A model trained on 2020 election news performs poorly on 2024 climate misinformation. This is called distribution shift, and it's the fundamental limitation of static ML classifiers.

Mitigation Strategies

  • 1.Regular retraining — Update the model monthly with fresh labeled data
  • 2.Feature monitoring — Track which features are drifting from training distribution
  • 3.Ensemble methods — Combine multiple models trained on different time periods
  • 4.Human-in-the-loop — Flag low-confidence predictions for human review

Ethical Considerations

Building a fake news detector raises important questions:

  • Who decides what's "fake"? — The ground truth labels in training data encode someone's judgment
  • Bias amplification — Models can disproportionately flag content from certain political perspectives
  • Censorship risk — Automated systems shouldn't be the sole arbiter of truth
  • Adversarial attacks — Bad actors can intentionally craft text to fool the classifier

The responsible approach is to use these systems as assistive tools — flagging content for human review rather than automatic removal.


Code: GitHub

NLPFake NewsPythonML
Suyash

Suyash Vakhariya

AI Engineer & Technical Product Manager. Building production AI systems.

© 2026 Suyash Vakhariya