Audio │ file-based
AI audio detection, with waveform playback.
Drop in an audio file. Detecto returns an aggregate synthesis-likelihood, a voice-clone heuristic, and a re-encode trace beside a waveform scrubber for review.
Try a scan
Drop a file here, or pick one
This is a quick demo, files are not saved. Sign in to keep your reports.
1 credit per 10 seconds
No account needed. Anonymous scans are rate-limited per IP.
Built for the people who review other people’s content
- Editors at independent magazines
- University writing centers
- Newsroom verification desks
- PR review teams
- Content QA leads at agencies
File-based
Accepted inputs
Audio files, ≤ 50 MB.
MP3, WAV, AAC, M4A. Designed for clips up to ~10 minutes.
Field interviews
On-the-record interviews submitted from the field for verification.
Spokesperson audio
Sound bites and audio statements supplied by client-facing representatives.
Podcast inserts
Insert clips proposed for inclusion in podcasts or audio editorials.
Verification escalations
Audio flagged for a second look during the verification cycle.
What we measure
Signals exposed in every audio report.
Aggregate
Synthesis likelihood
Aggregated likelihood across the synthesis-indicator ensemble.
Playback
Waveform scrubber
Audio playback with a waveform scrubber, shown beside the aggregate synthesis signal.
Voice
Voice-clone heuristic
Heuristic flag for cloned voices; not a substitute for speaker authentication.
Spectral
Spectral residual
Residual patterns at high frequency that recur in common synthesis pipelines.
Metadata
Re-encode trace
Markers of multiple encode passes; a clean recording is unlikely to show heavy re-encode trace.
Calibration
Confidence variance
Variance across resampled detection passes on the same audio file.
Honest limitations on audio detection
A detection score is probabilistic, not proof. Treat the report as evidence to discuss, not a verdict to publish.
- Heavily compressed audio (e.g., 96 kbps social re-shares) erodes synthesis indicators.
- Voice-clone heuristics are rough and not a substitute for speaker authentication.
- Mixed-source audio (real interview with synthesized voiceover) requires human review of the playback; the aggregate score is the wrong number alone.
- Background noise can mask synthesis artifacts; clean audio gives the detector more signal.
- No URL ingestion or live-call callbacks in v1 │ file-based only.
For PR teams and newsrooms
Get the synthesis indicator before the segment runs.
Your report is saved privately to your account, never public unless you share it.