According to recent employment fraud statistics, fake job postings have skyrocketed, costing applicants millions of dollars and countless hours of leaked personal data. Attackers no longer just copy-paste poorly written emails; they use generative AI to draft pristine, convincing descriptions, launch highly targeted phishing links, and build fake corporate personas.
Frustrated by this trend, I decided to build a solution: ScamShield, a real-time, AI-powered job scam detector that audits listings, validates URL safety, checks company registrations, and analyzes recruiter signals to calculate a real-time trust score before you ever hit “Apply.”

Here is an inside look at how it works, the architectural upgrades that boosted its accuracy to 98.11%, and why a multi-pillar approach is necessary to defeat modern scams.
The Core Philosophy: Why Text Analysis Alone Fails

When developers attempt to build scam detectors, they often make the mistake of relying solely on an NLP text classifier. If a job description looks normal, the system passes it.
But modern scammers are smart. They use corporate templates. The text might look flawless, but the application link points to a malicious domain, or the recruiter forces communication through an unverified personal email channel.
To stop this, ScamShield evaluates job postings through a 7-Stage Horizontal Verification Pipeline built on 5 Trust Pillars

1. NLP Text Analysis (The Brain)
The platform evaluates the semantics of the listing using a machine learning text classifier while auditing vocabulary density for clickbait signals. Crucially, it tracks lexical complexity via the Flesch Reading Ease index to spot overly convoluted, machine-generated text templates.
2. URL Phishing Scan (The Link)
It automatically inspects embedded links for secure SSL protocols (https:// vs. insecure http://) and audits domains against suspicious free hosting providers (like WordPress, Blogspot, or Wix) and URL shorteners commonly used to mask malicious sites.
3. Company Check (The Magnifying Glass)
ScamShield looks for description completeness and audits company names against official corporate registry structures, verifying standard legal suffixes such as Ltd, Inc, LLP, or Pvt.
4. Recruiter Behavior Check (The Shield)
This pillar flags high-risk signals, including personal contact details embedded in the body text (like random Gmail or Yahoo addresses instead of corporate domains), high-pressure urgency terms (e.g., “URGENT”, “IMMEDIATE”), and locations tied to historically high regional fraud ratios.
5. Trust Score Engine (The Metric)
By integrating machine learning probabilities with heuristic risk penalties, the system computes a final credibility rating from 0 to 100%:
- 80% - 100%: High Trust (Verified Safe / Approved)
- 50% - 79%: Neutral Risk (Verification Advised)
- Below 50%: Low Trust (Scam Alert / Application Blocked)
The 98% Accuracy: Under the Hood

To make sure the platform didn’t constantly flag legitimate jobs (avoiding “false alarms”), I recently completely overhauled the machine learning pipeline. The upgrades made a massive impact on performance:
- Upgrading the Text Classifier: I moved from a basic
CountVectorizerto aTfidfVectorizer(with 10,000 max features) and trained anSGDClassifier(loss='log_loss'). By handling data imbalance using SMOTE, the individual text F1-score jumped from 76.52% to 80.97%.
- Replacing Linear Models for Tabular Data: Linear models failed to capture non-linear correlations like the combination of character lengths, location threat ratios, and readability scores. Swapping this out for a tree-based
RandomForestClassifier(100 estimators) caused the tabular F1-score to rocket from a weak 16.54% to 56.32%.
- Ensemble Soft Voting: By combining the outputs using a weighted soft-voting model (0.5 {Text Risk} + 0.5 {Tabular Risk}), the final system achieved an incredible 98.11% Ensemble Accuracy and slashed false alarm flags by 90%.
A Premium User Experience
A robust backend means nothing if the frontend isn’t intuitive for a job seeker. Built on a lightweight, modular Flask architecture, the user interface features:
- An Animated Pipeline Tracker: Glowing neon tracks light up node-by-node during a live scan.
- A Trust Pillars Grid: Clean, hover-responsive cards display status badges (Passed in green, Warning in amber, Failed in red) for all four checks.
- A Seamless Theme Switcher: Users can easily toggle between an elegant Obsidian Dark Theme and a clean Alabaster Light Theme.
- Smart URL Scraping: Instead of manually copying and pasting details, users can simply drop a LinkedIn or Indeed link into the Smart URL Scan tab. The system’s crawler automatically extracts the job title, company name, location, and description, pre-populating fields for instantaneous threat analysis.
Open Source & Next Steps

This project was built to give power back to job seekers. The full source code, data pipelines, and setup guides are entirely open-source.
If you are an engineer, security enthusiast, or job seeker looking to explore the code or contribute to the pipeline, check out the repository here:
👉 Explore ScamShield on GitHub
Have you encountered a job scam during your search? What red flags do you look out for? Let’s discuss in the comments!

