A forensic platform that inspects video, audio, and images for signs of manipulation. Rather than leaning on a single black-box frame classifier, the video path extracts physiological signals - a pulse waveform recovered from subtle skin-colour changes (rPPG), motion from optical flow, and micro-expression strain - then maps them onto a simulated quantum state to measure how far they drift from calibrated biological baselines.
Synthetic media detectors often reduce a clip to per-frame spatial features and let a neural network decide. That approach is opaque and can generalise poorly to new generators. This project explores the opposite direction: read the physical and physiological traces that a genuine recording carries and that a synthesiser has no reason to reproduce faithfully.
The objective was a single application that unifies three analysis paths - video, audio, and still image - and returns a REAL / DEEPFAKE verdict with supporting metrics. The canonical implementation lives in 00_Unified_Deepfake_System; an older video-only prototype is retained under 02_Video_Deepfake_1 for archival purposes only.
Scope note
The "quantum" component is a classical numerical simulation of quantum dynamics running on a CPU. It is a mathematical framework for measuring signal complexity - not execution on quantum hardware, and no quantum-advantage claim is made.
System architecture
The component map below follows the video path, which is the most involved of the three. The audio and image paths share the same interface but substitute their own detection stages.
Video-path component map. The audio path replaces the last three blocks with librosa MFCC features and a scikit-learn classifier (audio_classifier.pkl + scaler.pkl); the image path replaces them with OpenCV error-level analysis. Only the video path uses the simulated quantum stage.
How data moves through the system
Capture or upload
A video, image, or audio clip is submitted through the Flask web UI, the command-line experiment runner, or the optional Chrome extension, which forwards browser frames to the local backend on port 8080.
Implemented
Face and landmark detection
MediaPipe FaceMesh locates the face across frames. Footage without enough usable face samples is reported as insufficient input rather than forced into a verdict.
Implemented
Physiological signal extraction
The rPPG stage recovers a pulse waveform from subtle skin-colour variation, Farneback optical flow estimates motion energy, and a micro-expression stage measures neuromuscular strain around the face.
Implemented
Normalisation and feature packaging
The extracted time-series are normalised and packaged into the feature vector consumed by the decision stages.
Implemented
Quantum-state simulation
Features are encoded into a simulated quantum state (PennyLane default.mixed) so that observables such as purity, coherence, and entanglement entropy can be read out as a complexity signal.
Simulated
Threshold-based verdict
The measured metrics are compared against calibrated biological boundaries. When at least two signals violate their thresholds, the clip is flagged as a deepfake.
Implemented
Parallel audio and image paths
Audio is classified from librosa MFCC and delta features through a joblib model and scaler; still images are scored with OpenCV error-level analysis using a deterministic anomaly score.
Implemented
Verified technology stack
Confirmed from the repository, its requirement pins, and the bundled model files. No video deep-learning model is shipped - the video decision is rule-based.
Language and runtimePython 3.11 in a virtual environment
Web and APIFlask with Flask-SocketIO for upload and live progress
Signal processingNumPy and SciPy for filtering and feature math
Audio MLlibrosa MFCC + delta features; scikit-learn classifier and scaler stored as audio_classifier.pkl and scaler.pkl
Quantum simulationPennyLane (qml.device("default.mixed")) running on CPU
Decision datacalibrated_thresholds.json - biological boundaries for the video verdict
Key technical decisions and trade-offs
Rule-based video decision over a learned model
The repository contains no .pt, .pth, or .onnx model for frame classification. Decisions come from explicit physiological boundaries in calibrated_thresholds.json. This keeps the verdict inspectable and needs no labelled training set, at the cost of sensitivity to lighting, makeup, and framerate.
Quantum stage as a measuring instrument
The simulated state is used to express purity, coherence, and entanglement entropy of the biological signals, giving a compact complexity reading. Keeping it on CPU keeps the system portable; it also caps throughput and is explicitly not a hardware speedup.
Fail honestly on unusable input
When a clip yields too few face-detected samples for bandpass filtering, or is too short for audio classification, the tool reports insufficient data instead of inventing a confidence value.
Three modalities with graceful degradation
Audio classification loads independently, and image scoring is heuristic, so the application still returns a useful result when no trustworthy biological signal is available.
Challenges and limitations
Threshold dependency. Accuracy is heavily dependent on the fixed biological thresholds, which may fail on extreme lighting, heavy makeup, or non-standard framerates.
CPU-bound simulation. The quantum component runs on classical hardware, limiting real-time throughput.
Input requirements. Audio needs clean clips of sufficient duration; very short clips are rejected. Video needs a sufficiently visible face.
Backend coupling. The Chrome extension only works while the local Flask backend is running on port 8080.
No broad benchmarking yet. The repository is a functional prototype; generalisation across diverse, in-the-wild datasets still needs empirical validation, so no accuracy figures are claimed here.
Legacy code. The 02_Video_Deepfake_1 prototype is archival and is not the primary runtime.
Reproducing and exploring it
The repository documents a Python 3.11 virtual environment and a single Flask entry point. MediaPipe must be installed exactly as pinned because the FaceMesh pipeline depends on it.
git clone https://github.com/mananbd07/Deepfake-System.git
cd Deepfake-System
python -m venv venv
venv\Scripts\activate
pip install -r requirements.txt
cd 00_Unified_Deepfake_System
python qbioforge_web/app.py
Then open http://localhost:8080. A clip can also be run headlessly with python -m experiments.run_single_video --video <path_to_mp4>. Test media is not committed - datasets and sample videos are excluded by .gitignore, so you supply your own.
Future engineering work
Future work - not implemented
Benchmark the pipeline against public deepfake datasets to substantiate generalisation.
Calibrate or learn the physiological thresholds automatically while keeping the decision interpretable.
Broaden capture beyond the local extension and webcam inputs.
Add efficiency work so the simulated quantum stage keeps up with longer clips.
Related case studies
This pipeline shares its signal-extraction approach with the packaged QBioForge application, and its applied-vision style with the traffic system.