Case study

Unified Deepfake Detection

A forensic platform that inspects video, audio, and images for signs of manipulation. Rather than leaning on a single black-box frame classifier, the video path extracts physiological signals - a pulse waveform recovered from subtle skin-colour changes (rPPG), motion from optical flow, and micro-expression strain - then maps them onto a simulated quantum state to measure how far they drift from calibrated biological baselines.

Overview and objective

Synthetic media detectors often reduce a clip to per-frame spatial features and let a neural network decide. That approach is opaque and can generalise poorly to new generators. This project explores the opposite direction: read the physical and physiological traces that a genuine recording carries and that a synthesiser has no reason to reproduce faithfully.

The objective was a single application that unifies three analysis paths - video, audio, and still image - and returns a REAL / DEEPFAKE verdict with supporting metrics. The canonical implementation lives in 00_Unified_Deepfake_System; an older video-only prototype is retained under 02_Video_Deepfake_1 for archival purposes only.

Scope note

The "quantum" component is a classical numerical simulation of quantum dynamics running on a CPU. It is a mathematical framework for measuring signal complexity - not execution on quantum hardware, and no quantum-advantage claim is made.

System architecture

The component map below follows the video path, which is the most involved of the three. The audio and image paths share the same interface but substitute their own detection stages.

Component map of the Unified Deepfake Detection video pipeline A four-stage stack: the Flask input interface feeds biological signal extraction, which feeds a simulated quantum-state stage, which feeds a threshold-based decision. Input and interface Flask + SocketIO web app (port 8080) or CLI: experiments/run_single_video Biological signal extraction MediaPipe FaceMesh facial landmarks rPPG (POS), optical flow, micro-strain Quantum-state simulation PennyLane default.mixed, CPU only simulated state - not quantum hardware Threshold decision calibrated_thresholds.json boundaries two or more violations flag the clip
Video-path component map. The audio path replaces the last three blocks with librosa MFCC features and a scikit-learn classifier (audio_classifier.pkl + scaler.pkl); the image path replaces them with OpenCV error-level analysis. Only the video path uses the simulated quantum stage.

How data moves through the system

  1. Capture or upload

    A video, image, or audio clip is submitted through the Flask web UI, the command-line experiment runner, or the optional Chrome extension, which forwards browser frames to the local backend on port 8080.

    Implemented
  2. Face and landmark detection

    MediaPipe FaceMesh locates the face across frames. Footage without enough usable face samples is reported as insufficient input rather than forced into a verdict.

    Implemented
  3. Physiological signal extraction

    The rPPG stage recovers a pulse waveform from subtle skin-colour variation, Farneback optical flow estimates motion energy, and a micro-expression stage measures neuromuscular strain around the face.

    Implemented
  4. Normalisation and feature packaging

    The extracted time-series are normalised and packaged into the feature vector consumed by the decision stages.

    Implemented
  5. Quantum-state simulation

    Features are encoded into a simulated quantum state (PennyLane default.mixed) so that observables such as purity, coherence, and entanglement entropy can be read out as a complexity signal.

    Simulated
  6. Threshold-based verdict

    The measured metrics are compared against calibrated biological boundaries. When at least two signals violate their thresholds, the clip is flagged as a deepfake.

    Implemented
  7. Parallel audio and image paths

    Audio is classified from librosa MFCC and delta features through a joblib model and scaler; still images are scored with OpenCV error-level analysis using a deterministic anomaly score.

    Implemented

Verified technology stack

Confirmed from the repository, its requirement pins, and the bundled model files. No video deep-learning model is shipped - the video decision is rule-based.

  • Language and runtimePython 3.11 in a virtual environment
  • Web and APIFlask with Flask-SocketIO for upload and live progress
  • Computer visionMediaPipe FaceMesh (pinned mediapipe==0.10.9), OpenCV (opencv-python-headless)
  • Signal processingNumPy and SciPy for filtering and feature math
  • Audio MLlibrosa MFCC + delta features; scikit-learn classifier and scaler stored as audio_classifier.pkl and scaler.pkl
  • Quantum simulationPennyLane (qml.device("default.mixed")) running on CPU
  • Decision datacalibrated_thresholds.json - biological boundaries for the video verdict

Key technical decisions and trade-offs

Rule-based video decision over a learned model

The repository contains no .pt, .pth, or .onnx model for frame classification. Decisions come from explicit physiological boundaries in calibrated_thresholds.json. This keeps the verdict inspectable and needs no labelled training set, at the cost of sensitivity to lighting, makeup, and framerate.

Quantum stage as a measuring instrument

The simulated state is used to express purity, coherence, and entanglement entropy of the biological signals, giving a compact complexity reading. Keeping it on CPU keeps the system portable; it also caps throughput and is explicitly not a hardware speedup.

Fail honestly on unusable input

When a clip yields too few face-detected samples for bandpass filtering, or is too short for audio classification, the tool reports insufficient data instead of inventing a confidence value.

Three modalities with graceful degradation

Audio classification loads independently, and image scoring is heuristic, so the application still returns a useful result when no trustworthy biological signal is available.

Challenges and limitations

  • Threshold dependency. Accuracy is heavily dependent on the fixed biological thresholds, which may fail on extreme lighting, heavy makeup, or non-standard framerates.
  • CPU-bound simulation. The quantum component runs on classical hardware, limiting real-time throughput.
  • Input requirements. Audio needs clean clips of sufficient duration; very short clips are rejected. Video needs a sufficiently visible face.
  • Backend coupling. The Chrome extension only works while the local Flask backend is running on port 8080.
  • No broad benchmarking yet. The repository is a functional prototype; generalisation across diverse, in-the-wild datasets still needs empirical validation, so no accuracy figures are claimed here.
  • Legacy code. The 02_Video_Deepfake_1 prototype is archival and is not the primary runtime.

Reproducing and exploring it

The repository documents a Python 3.11 virtual environment and a single Flask entry point. MediaPipe must be installed exactly as pinned because the FaceMesh pipeline depends on it.

git clone https://github.com/mananbd07/Deepfake-System.git
cd Deepfake-System
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
cd 00_Unified_Deepfake_System
python qbioforge_web/app.py

Then open http://localhost:8080. A clip can also be run headlessly with python -m experiments.run_single_video --video <path_to_mp4>. Test media is not committed - datasets and sample videos are excluded by .gitignore, so you supply your own.

Future engineering work

Future work - not implemented
  • Benchmark the pipeline against public deepfake datasets to substantiate generalisation.
  • Calibrate or learn the physiological thresholds automatically while keeping the decision interpretable.
  • Broaden capture beyond the local extension and webcam inputs.
  • Add efficiency work so the simulated quantum stage keeps up with longer clips.

This pipeline shares its signal-extraction approach with the packaged QBioForge application, and its applied-vision style with the traffic system.