Detecto
Audio │ file-based

AI audio detection, with waveform playback.

Drop in an audio file. Detecto returns an aggregate synthesis-likelihood, a voice-clone heuristic, and a re-encode trace beside a waveform scrubber for review.

Try a scan

Drop a file here, or pick one

This is a quick demo, files are not saved. Sign in to keep your reports.

1 credit per 10 seconds

No account needed. Anonymous scans are rate-limited per IP.

Built for the people who review other people’s content

  • Editors at independent magazines
  • University writing centers
  • Newsroom verification desks
  • PR review teams
  • Content QA leads at agencies
File-based
Accepted inputs

Audio files, ≤ 50 MB.

MP3, WAV, AAC, M4A. Designed for clips up to ~10 minutes.

Field interviews
On-the-record interviews submitted from the field for verification.
Spokesperson audio
Sound bites and audio statements supplied by client-facing representatives.
Podcast inserts
Insert clips proposed for inclusion in podcasts or audio editorials.
Verification escalations
Audio flagged for a second look during the verification cycle.
What we measure

Signals exposed in every audio report.

Aggregate
Synthesis likelihood
Aggregated likelihood across the synthesis-indicator ensemble.
Playback
Waveform scrubber
Audio playback with a waveform scrubber, shown beside the aggregate synthesis signal.
Voice
Voice-clone heuristic
Heuristic flag for cloned voices; not a substitute for speaker authentication.
Spectral
Spectral residual
Residual patterns at high frequency that recur in common synthesis pipelines.
Metadata
Re-encode trace
Markers of multiple encode passes; a clean recording is unlikely to show heavy re-encode trace.
Calibration
Confidence variance
Variance across resampled detection passes on the same audio file.

Honest limitations on audio detection

A detection score is probabilistic, not proof. Treat the report as evidence to discuss, not a verdict to publish.

  • Heavily compressed audio (e.g., 96 kbps social re-shares) erodes synthesis indicators.
  • Voice-clone heuristics are rough and not a substitute for speaker authentication.
  • Mixed-source audio (real interview with synthesized voiceover) requires human review of the playback; the aggregate score is the wrong number alone.
  • Background noise can mask synthesis artifacts; clean audio gives the detector more signal.
  • No URL ingestion or live-call callbacks in v1 │ file-based only.
For PR teams and newsrooms

Get the synthesis indicator before the segment runs.

Your report is saved privately to your account, never public unless you share it.