From Your Microphone to the Other Screen: The Journey of a Video Call
The moment you click "Join," your device's camera and microphone begin capturing a continuous stream of raw data. That raw feed is immediately compressed by a codec — a piece of software that encodes (compresses) outgoing data and decodes incoming data. Codecs like H.264, VP9, and AV1 for video, and Opus for audio, strip away information the human eye and ear are unlikely to notice, reducing file sizes dramatically without obvious quality loss.
This compressed stream is then packaged into small data units and transmitted across the internet. How it travels — and who can see it along the way — depends on the platform's architecture.
~50–100 kbps
Typical audio data rate per call participant
Standard voice codecs used in video calling compress audio to roughly 50–100 kilobits per second, a fraction of raw microphone output.
1–4 Mbps
Typical HD video call bandwidth per stream
HD video in a video call generally requires between 1 and 4 megabits per second depending on resolution, frame rate, and codec efficiency.
~60–80%
Data reduction achieved by modern video codecs
Modern codecs like H.265 and AV1 can reduce video file sizes by 60–80% compared to uncompressed video, enabling practical real-time transmission.
Server-Routed vs. Peer-to-Peer: Who Sits in the Middle?
There are two broad models for how call data moves between participants:
- Peer-to-peer (P2P): Data flows directly between participants' devices, bypassing the platform's servers. This limits the platform's technical exposure to your call content. However, P2P requires both devices to share their IP addresses with each other, and it becomes less practical for group calls with many participants.
- Server-routed (relay): Data passes through the platform's infrastructure before reaching the other participant. This is the standard approach for group video calls and offers better reliability across different network conditions — but it means the platform's servers handle your audio and video in transit.
Whether that server access is meaningful depends on encryption. For a deeper look at how encryption protects — and doesn't protect — data in transit, see our explainer on what encryption actually protects in messaging apps.
Metadata Is Collected Even Without Recording
Even platforms that never record call content still collect metadata — participant identifiers, call duration, device type, and approximate location. Metadata can reveal communication patterns without capturing any spoken words. This is standard practice across virtually all consumer video calling services.
Encryption: The Key Variable in Call Privacy
Encryption determines whether a platform that routes your call through its servers can actually read your audio and video data. There are two main approaches:
- Transport encryption (TLS/DTLS): Data is encrypted between your device and the platform's server, but the platform can decrypt it on the server side. This protects against outside eavesdroppers but not against the platform itself.
- End-to-end encryption (E2EE): Data is encrypted on your device and can only be decrypted by the recipient's device. Even if it passes through the platform's servers, those servers see only unreadable data.
E2EE for video calls is technically more demanding than for text messages and is not universally implemented. Some platforms offer it only in one-on-one calls, or require users to manually enable it. Always verify the encryption status in a platform's documentation rather than assuming it is active.
Check Encryption Status Before Sensitive Calls
Before a call involving sensitive personal, medical, or business information, verify whether the platform activates end-to-end encryption by default or only on request. This information is usually available in the platform's security documentation or help center — not just marketing copy.
AI Features and What They Mean for Your Audio
Many modern video calling apps include AI-powered features: background noise cancellation, automatic transcription, live captions, or meeting summaries. These are genuinely useful — but they raise a straightforward question: where does the AI processing happen?
If processing is on-device, your audio is analyzed locally and never leaves your machine. If processing is cloud-based, audio samples are sent to the platform's servers for analysis. Cloud processing is common because it allows more powerful models to run without taxing your device, but it means audio data travels beyond your device in ways a basic call would not require.
This is similar to the dynamic described in our article on what smart speakers actually capture and when — the distinction between local and cloud processing matters significantly for understanding your data exposure.
What Platforms Typically Retain After the Call
Even when a platform doesn't record call content, it typically retains metadata: who called whom, when, for how long, from which device type, and from which approximate location. This metadata can reveal patterns about your communication habits even without capturing a single word spoken.
Call recordings, when enabled, are stored according to the platform's retention schedule — which varies widely. Some delete recordings automatically after a set period; others retain them until the user manually deletes them. AI-generated transcripts or summaries, where offered, are usually stored separately from any raw audio.
The most reliable source of information on any specific platform's practices is its own privacy policy and data processing documentation. These documents can be dense, but they are legally binding and must disclose what data is collected, how it is used, and for how long it is retained. For a broader comparison of how communication apps differ in their data practices, our overview of key differences between messaging apps provides useful context.
Frequently Asked Questions
Platforms generally require consent to record calls, and most consumer apps notify all participants when recording starts. However, policies vary — always read the platform's terms of service and privacy policy. Some platforms reserve the right to collect metadata or short audio samples to improve AI features.
End-to-end encryption (E2EE) means your audio and video data is encrypted on your device and can only be decrypted by the intended recipient — not the platform's servers. Not all video calling apps offer E2EE by default, and some only apply it in specific call modes.
For most mainstream apps, yes — call data is routed through the company's servers, especially when calls involve multiple participants. Peer-to-peer connections, which bypass servers entirely, are used by some apps under specific conditions but are less common for group calls.
They can be, depending on implementation. If AI processing happens locally on your device, your audio never leaves it. If processing happens in the cloud, audio samples may pass through the provider's servers. Check whether the feature is described as on-device or cloud-based.
Most platforms retain metadata — call duration, participant identifiers, timestamps — even if they don't store call content. Some platforms retain recordings if the user activates that feature. Reviewing the platform's privacy policy is the most reliable way to understand retention practices.
The content on this site is for informational purposes only and is not a substitute for professional advice. Always consult a qualified professional for guidance specific to your situation.

