Deep dive
Build video into a workflow
General video meetings are a mature, crowded market with free options everywhere. The opportunity is video inside a workflow: a doctor's consultation with notes and prescriptions, a tutoring session with a whiteboard and homework, a hiring interview with scorecards, or a sales demo logged straight into the CRM. Each removes steps that a general tool leaves to the user.
That focus also simplifies the MVP. A clinic needs reliable one-to-one calls, waiting rooms and consent, not webinars for thousands. A school needs classroom controls and recordings, not dial-in numbers. Our product discovery work turns the workflow into a backlog and decides which meeting features actually matter.
How multi-party video works
Each participant's app opens a WebRTC connection to a selective forwarding unit, or SFU. Every sender uploads one stream, usually in several qualities through simulcast; the SFU forwards to each receiver the quality that suits their bandwidth and window size. Unlike older mixing servers, an SFU does not decode video, so one server can handle many meetings at low latency.
Signalling, which sets up those connections, runs over WebSockets to your backend. TURN relays carry media when firewalls block direct paths, which is common on corporate networks. Measure round-trip time, packet loss, jitter and freezes per call, and route each meeting to the media region closest to most participants.
- Use simulcast or SVC so one slow participant does not degrade everyone else in the meeting.
- Prioritise audio over video when bandwidth drops; people forgive a frozen frame, not lost words.
- Run TURN over TLS on port 443 for networks that block everything else.
Recording, transcripts and AI summaries
Cloud recording joins the meeting as a hidden participant, composes the streams into a layout and encodes a video file. Transcription runs on the audio, ideally with speaker separation, and a language model produces summaries, decisions and action items. Store transcripts as searchable text so users can find what was said across months of meetings.
Accuracy and trust matter more than features. Show summaries with links back to the transcript, let hosts edit and delete them, and be clear about where audio is processed. If customers need data to stay in a region, use speech and language models that can run in that region, including open models hosted in your own cloud. Our guide to AI chatbot and agent costs covers how inference pricing adds up.
Security, privacy and consent
All WebRTC media is encrypted in transit by default, but the SFU can access it. End-to-end encryption with SFrame, standardised as RFC 9605, encrypts media on the sender's device so media servers only route it. It limits cloud recording and server-side AI, so most products offer it as an option for sensitive meetings rather than the default.
Recording laws differ: some jurisdictions require every participant's consent, so show clear indicators and notices. Healthcare calls need HIPAA safeguards in the US, and personal data in recordings falls under the GDPR and India's DPDP Act. If you show AI-generated summaries to users in the EU, the AI Act's transparency rules, in force since August 2026, are worth reviewing with counsel.
Running costs and scaling
Video is a usage business. Media server compute, TURN bandwidth, recording storage and processing and AI inference all scale with meeting minutes, not with users. Model cost per participant hour before setting free-tier limits, and watch TURN usage closely, since relayed media is the most expensive path.
Scaling means more regions and more automation: autoscaling media servers by active meetings, draining servers gracefully before updates and monitoring call quality as a product metric. Maintenance is roughly 15-20% of the build cost per year on top of infrastructure. Our DevOps consulting team usually owns this capacity planning.