JSONMeet. Slack, Meet i TeamSpeak w jednym, na własnym serwerze.
JSONMeet. Slack, Meet i TeamSpeak
w jednym, na własnym serwerze.
- Zespół rozbity na 4 narzędzia: głos, czat, spotkania, nagrania
- Nagrania rozmów sprzedażowych w cudzej chmurze
- Klient ma wejść z linku, bez konta i instalacji
- AI ma pomagać w trakcie rozmowy, nie po niej
- Kanały głosowe i tekstowe, DM-y, wyszukiwarka
- Spotkania pod linkiem z poczekalnią i nagrywaniem
- Prezenter HQ bez rekompresji na serwerze
- Sufler AI: transkrypcja na żywo i podpowiedzi
- Jeden serwer w Warszawie, ~10 ms pingu z Polski
- Zero abonamentów per użytkownik
- Recordings and transcripts u nas, z retencją
- Aplikacja desktopowa z globalnym push-to-talk
Cztery narzędzia, cztery abonamenty, zero kontroli nad nagraniami.
Nasz zespół pracuje zdalnie, a rozmowy handlowe traktujemy jak dane, których nie oddajemy w cudze ręce. Przez cały dzień siedzieliśmy na kanale głosowym w TeamSpeaku, pisaliśmy na Slacku i Discordzie, a spotkania z klientami umawialiśmy w Google Meet. Każde narzędzie robiło swoją część dobrze. Problem był w sumie.
Cztery loginy, cztery zestawy powiadomień, cztery miejsca, w których ginęły ustalenia. Nagrania rozmów sprzedażowych lądowały w chmurze dostawcy, na jego warunkach i z jego retencją. Chcieliśmy je transkrybować, analizować i uczyć się na nich. Każda taka próba oznaczała eksport, upload do kolejnego narzędzia i kolejną umowę powierzenia.
Największa luka była w samej rozmowie. Handlowiec słyszy obiekcję klienta i ma kilka sekund na odpowiedź. Cała wiedza firmy o tej obiekcji (co zadziałało, czego unikać, jaki jest kontekst leada z CRM) leży w dokumentach, do których w trakcie rozmowy nikt nie sięga. AI podsumowujące spotkanie po fakcie nic tu nie zmienia.
Postawiliśmy wymagania: jedna aplikacja, na naszym serwerze, klient wchodzi z linku bez konta, jakość udostępniania ekranu wystarczająca do prezentacji interfejsów w 4K, nagrywanie z wyraźnym banerem dla wszystkich uczestników. I sufler, który słucha rozmowy i podpowiada handlowcowi na żywo, niewidoczny dla klienta.
Co zbudowaliśmy. Cztery moduły JSONMeet.
Komunikator: kanały, czat, DM-y
Układ jak w Discordzie: workspace’y, kanały tekstowe i głosowe w kategoriach, panel członków ze statusami na żywo. Trwały czat z wzmiankami, plikami, reakcjami i odpowiedziami. Wiadomości prywatne. Pełnotekstowa wyszukiwarka po kanałach i własnych rozmowach. Wszystko po jednym połączeniu WebSocket: wiadomości, obecność, „pisze…”, liczniki nieprzeczytanych.
Spotkania z klientem i nagrywanie
Pokój pod linkiem. Członek zespołu wchodzi od razu, gość klienta trafia do poczekalni i czeka na wpuszczenie jednym kliknięciem. Gość nie dotyka serwera mediów, dopóki gospodarz nie zatwierdzi. Nagrywanie po stronie serwera do MP4 z banerem „Nagrywanie” widocznym dla wszystkich, lista nagrań z retencją i automatycznym kasowaniem.
Prezenter HQ i udostępnianie ekranu do 4K
W przeglądarce presety 1080p60, 1440p60, 1440p30 „ostry tekst” i 4K30, zmieniane w trakcie. Do prezentacji studyjnej osobna aplikacja desktopowa: OBS lub GStreamer koduje obraz sprzętowo i wysyła go protokołem WHIP prosto do serwera mediów, który przekazuje strumień dalej bez rekompresji. Serwer prawie nie zużywa CPU, a odbiorcy widzą to, co wyszło z karty graficznej prezentera.
Sufler AI dla handlowca
Transkrypcja rozmowy na żywo trafia do modelu językowego z playbookiem firmy: macierz 26 obiekcji, zasady negocjacji cenowych, brief o firmie. Sufler pokazuje handlowcowi wykrytą obiekcję, gotową odpowiedź, pytania do zadania i wskaźnik gotowości do domknięcia w skali 0 do 100. Przed rozmową wybierasz leada z CRM i sufler dostaje jego kontekst. Klient tego panelu nigdy nie widzi. Po rozmowie: analiza, scorecard, wyszukiwanie po transkrypcjach.
Pod maską
Frontend w React 18 z Vite i Tailwind CSS 4. Backend w Node.js na Fastify 5 z Prismą 6 i PostgreSQL 17, Redis 7 do sesji, obecności i kolejek. Media obsługuje serwer LiveKit w wersji open source, postawiony samodzielnie: SFU dla dźwięku i wideo, Egress do nagrywania, Ingress do przyjmowania strumienia WHIP z prezentera. Kamera idzie w VP9 z trzema warstwami jakości, ekran w H.264 simulcast, więc słabszy odbiorca dostaje lżejszą warstwę bez pogarszania obrazu innym.
Redukcja szumów (RNNoise) i wirtualne tło (MediaPipe) działają lokalnie w przeglądarce uczestnika. Za firmowymi firewallami ratuje wbudowany TURN po TLS i awaryjny transport po TCP. Aplikacja desktopowa w Electronie łączy się z instancją firmy i sama dostaje każdą aktualizację interfejsu; dodaje tray, globalny skrót push-to-talk, licznik nieprzeczytanych na ikonie i deep-linki do spotkań.
Transkrypcja i model językowy suflera działają na naszych własnych serwerach GPU, połączonych z serwerem aplikacji prywatną siecią. Produkcja to jeden VPS w OVHcloud w Warszawie (8 vCPU, 16 GB RAM) z usługami pod systemd, bez Dockera. Wdrożenie: archiwum z gita, build na serwerze, migracje, restart usług.
Przed i po. Ten sam zespół, jedno narzędzie.
- 4 narzędzia: TeamSpeak, Slack/Discord, Google Meet, osobne nagrywanie
- Nagrania rozmów in the provider's cloud, retention on its terms
- Klient instaluje lub loguje się, żeby wejść na spotkanie
- Wiedza o obiekcjach w dokumentach, poza rozmową
- Ekran udostępniany w jakości, jaką narzędzie uzna za stosowną
- Abonament per użytkownik w każdym z narzędzi
- 1 aplikacja: kanały głosowe, czat, DM-y, spotkania, nagrania
- Recordings and transcripts na własnym serwerze, retencja ustawiana per plan
- Klient wchodzi z linku w przeglądarce, przez poczekalnię
- Sufler podpowiada w trakcie rozmowy, klient go nie widzi
- Presety do 4K30 i 1440p60, prezenter bez rekompresji
- Jeden VPS, koszt stały niezależnie od liczby osób
Liczby projektu.
99 commitów
frontend, backend, desktop, prezenter
test: 10 nadawców < 50% CPU
plus negocjacje cenowe
Do tego: 22 migracje bazy, 26 plików testów backendu, 4 plany z egzekwowanymi limitami (FREE, STARTER, PRO, ENTERPRISE), 2 równoległe nagrania na serwerze, deep-linki jsonmeet://join/<kod> z aplikacji desktopowej.
Nie chodziło o to, żeby mieć własnego Meeta dla zasady. Chodziło o to, że rozmowa sprzedażowa to najcenniejsze dane w firmie, a my oddawaliśmy je komuś innemu. Teraz nagranie, transkrypcja i podpowiedzi są nasze. I działają w trakcie rozmowy, nie po niej.
Jak wyglądała budowa. Osiem tygodni.
1
Projekt i fundament
Dokument projektowy: wymagania, ryzyka, model danych, plan etapów. Fundament: wiele firm w izolacji na jednej instancji, konta i sesje, role z uprawnieniami rozwiązywanymi przy logowaniu, zaproszenia kopiowalnym linkiem. Serwer mediów postawiony na VPS z wbudowanym TURN.
2
Spotkania pod linkiem
Pokój z kodem, poczekalnia dla gości trzymana po stronie serwera, spotlight i siatka, udostępnianie ekranu z presetami. Nagrywanie całej kompozycji pokoju do MP4 z banerem dla uczestników i listą nagrań z retencją.
3
Prezenter HQ
Osobna aplikacja desktopowa dla prezentera: OBS lub GStreamer koduje sprzętowo, strumień idzie protokołem WHIP do Ingress i dalej bez rekompresji. Decyzja: WHIP zamiast RTMP, bo RTMP wymusza ponowne kodowanie na serwerze.
4–5
Komunikator
Stały układ aplikacji: pasek workspace’ów, kanały w kategoriach, treść, panel członków. Trwały czat z wzmiankami, plikami, reakcjami i odpowiedziami, DM-y, wyszukiwarka. Kanały głosowe działające niezależnie od nawigacji: klikasz po aplikacji, rozmowa trwa. Push-to-talk bez renegocjacji połączenia.
6
Aplikacja desktopowa
Electron łączący się z instancją firmy: tray, globalny skrót push-to-talk, licznik nieprzeczytanych, własny wybór okna do udostępnienia, deep-linki do spotkań. Pierwsze uruchomienie to adres serwera i akceptacja certyfikatu; aktualizacje interfejsu przychodzą z serwera.
7
Transkrypcja i sufler
Transkrypcja na żywo na własnym GPU. Sufler z playbookiem 26 obiekcji, briefem firmy i kontekstem leada pobieranym z CRM. Widoczny tylko dla zespołu, konfigurowany per workspace. Jedno wejście do modelu językowego dla wszystkich funkcji AI, więc zmiana modelu to zmiana konfiguracji.
8
Conversation analysis and production
Post-talk analysis with speaker recognition, scorecard, signals, and post-transcription search. Image and sound hygiene: noise reduction, virtual background, automatic lowering of the resolution when the USB camera drops to several frames. Implementation on one VPS in Warsaw under systemd, without Docker.
Self-hosted messaging for business: when your own server makes sense and when it doesn't
An article for business owners and sales leaders considering moving chat, voice calls, and client meetings onto their own infrastructure. We explain when a self-hosted business messenger is a sensible choice, what needs to be designed into it, and which risks to plan for from day one.
What a self-hosted messenger is and how it differs from Slack or Meet
A self-hosted communication platform is an app for chat, calls, and video conferencing running on a server the company manages (its own or a rented VPS). Messages, recordings, and transcripts stay in its own infrastructure, not in a SaaS vendor’s cloud.
The difference isn't just about where data is stored. With your own deployment, the company decides on recording retention, who has access to recordings, and which tools (e.g., AI models) may process them. With off-the-shelf services, the provider makes those decisions in its terms of service.
A self-hosted business messenger doesn't have to mean writing everything from scratch. A sensible implementation consists of mature open-source components (media server, database, queues) and a custom application layer that addresses the team's specific needs.
An alternative to Slack and Google Meet: when it's worth considering
For most companies, off-the-shelf tools are enough. A custom solution starts to make sense in three situations, which often occur together:
- client conversations are data you don't want to hand over (recordings, transcripts, sales notes),
- the team uses several tools in parallel and pays for each one per user,
- you want AI to work with your playbook and CRM data during the call, not just summarize it afterward.
In the implementation described above, the team worked with four tools: TeamSpeak for voice, Slack and Discord for chat, Google Meet for client meetings, and separate recording software. Four logins, four sets of notifications, and decisions got lost between them. JSONMeet replaced them with a single app.
| Issue | Off-the-shelf SaaS tools | Self-hosted messenger |
|---|---|---|
| Recordings and transcripts | in the provider's cloud, retention on its terms | on the company's server, with retention set by you |
| Billing | per-user subscription in every tool | server and maintenance cost, independent of headcount |
| AI during the call | limited to the vendor's features | any model, your own playbook, data from the CRM |
| Maintenance | on the vendor's side | handled by the company or a technical partner |
| Launch | immediately | design and implementation (here: 8 weeks) |
Recording calls on your own server: requirements worth writing down
Before the first line of code is written, it’s worth writing down the business requirements. In the project described, they were: one app, a self-hosted server, clients join from a link without an account or installation, screen sharing good enough to show an interface in 4K, and recording with a clear notice to all participants.
This leads to specific features. A guest lands in a waiting room and doesn't connect to the media server until the host lets them in. Recording happens server-side to an MP4 file, and a “Recording” banner is visible throughout the call. The recordings list has a retention policy with automatic deletion, configurable from 7 days to unlimited.
This last point matters for GDPR. If the recordings are stored with you, you don't need another data processing agreement every time you want to transcribe or analyze them with a different tool.
Architecture: WebRTC, media server, and client firewalls
The core of every messaging app with video is an SFU media server, which receives streams from participants and forwards them on. In JSONMeet, this is a self-hosted open-source LiveKit: SFU for audio and video, Egress for recording, and Ingress for receiving the stream from the presenter’s app.
The most common risk with self-hosting WebRTC is a corporate firewall on the client's side. It needs to be planned for from the start. Here, the server has built-in TURN over TLS on port 5349 and a fallback TCP transport, and for the most restrictive networks there's a second IP address with TURN on port 443, which looks like ordinary HTTPS traffic to the firewall.
The second decision is video quality. The camera goes out in VP9 with three quality layers and screen sharing in H.264 simulcast, so a recipient with a weaker connection gets a lighter layer without degrading the picture for others. For studio presentations, a separate app encodes the video in hardware and sends it via WHIP with no server-side re-encoding. Presets go up to 4K30 and 1440p60.
AI prompter for sales reps: AI during the call, not after it
A meeting summary after the fact doesn’t help much a sales rep who hears an objection and has a few seconds to respond. That’s why, in the implementation described, the call transcript is fed live to a language model together with the company’s playbook: a matrix of 26 objections, price negotiation rules, and a company brief.
Before the call, the sales rep selects a lead from the CRM, and the prompter receives its context. During the call, it shows the detected objection, a suggested response, questions to ask, and a readiness-to-close score on a scale of 0 to 100. The client doesn't see this panel. After the call, analysis, a scorecard, and transcript search are available.
An important detail: transcription and the model run on in-house GPU servers connected to the application server over a private network. Conversation content never goes to an external AI provider. A single entry point to the model for all AI features means that switching models is just a configuration change.
Maintaining a self-hosted messenger without a DevOps team
A common question is: who will maintain it? The answer depends on how simply you design production. JSONMeet runs on a single VPS in Warsaw (8 vCPU, 16 GB RAM) with services under systemd, without Docker or an orchestrator. Deployment is a single script: archive from git, build, database migrations, service restart.
In testing, this server handled a room with about 30 participants, 10 of whom were sharing video, at under 50% CPU load. Ping from Poland is about 10 ms. The Electron desktop app pulls interface updates from the server, so there's no need to update it manually on each computer.
For companies that don't want to run a server themselves, a sensible option is an instance maintained by a technology partner. The data still stays in dedicated infrastructure, and responsibility for updates and security is set out in the contract.
Before deciding, it's worth calculating three things: how many people and tools the change will affect, how valuable call recordings are to you, and whether AI should work on your data. If the answers point toward control over your data, a self-hosted messenger for your company is worth considering.
We design similar systems as part of custom web applications, and features based on language models, such as the live prompter, as part of AI implementations; we describe connecting the messenger to a CRM in ERP and CRM integrations.
Three questions we hear when we show JSONMeet.
Why own a messenger when Google Meet and Slack are ready and cheap?
For most companies, ready-made tools are enough. Own makes sense when customer conversations are data you don't want to give away (recordings, transcripts, sales notes), when you pay per user for four tools at once, or when you want an AI that runs on your playbook and your CRM during the conversation, not after it. JSONMeet was created for these three reasons at the same time.
WebRTC behind the client's company firewall will not work.
This is a real risk and we have planned it from day one. The media server has a built-in TLS TURN on port 5349 and emergency transport over TCP. In the scenario of the most restrictive client network, a second IP address is provided from the TURN on port 443, which looks like regular HTTPS traffic to the firewall. The guest enters from the browser, without installation.
Who will maintain it? We don't have a DevOps team.
Production includes one VPS and six services under the systemd: media server, API, reverse proxy, Redis and containers for recording and receiving the stream. The implementation is one script: archive from git, build, migrations, restart. There is no orchestrator, no Docker Compose. The desktop application updates the interface itself, from the server. For external customers, we offer JSONMeet as an instance maintained by us.
Does the AI souffler work during the conversation or only after it?
In progress. The transcript of the conversation flies live to the language model running on our GPUs, with a playbook of 26 objections and a lead context downloaded from the CRM. Sufler detects objections in the customer's statement, prompts ready answers and shows the readiness to close indicator on a scale of 0–100. The customer cannot see the airfoil panel. Analysis after the fact (scorecard, search after transcripts) is also available, but it does not replace the live souffler.
Are the call recordings left only on our server?
Yes, that was the whole game. The recording for MP4 goes to the volume connected to your VPS. Retention is configured per plan (from 7 days to no limit). The customer hears the message about the recording and sees the „Recording” banner throughout the conversation. Neither recordings nor transcripts leave your infrastructure — there is no external cloud provider.
How much does JSONMeet self-hosting cost compared to Slack + Meet + TeamSpeak fees?
It depends on the team, but the return on investment in JSONMeet typically starts with 8–15 people when the per-user costs in the finished tools start to add up. The specific bill depends on which modules you need — convert with the calculator below. We also offer an instance managed by us if you do not want your own DevOps.
Does JSONMeet replace CRM and sales tools?
No. JSONMeet is a communication layer that integrates with your CRM (Pipedrive, HubSpot, Salesforce, own API). The lead context from the CRM feeds the souvenir before the conversation, and the result of the conversation (transcription, scorecard, note) returns to the CRM after its completion. The sales call is in JSONMeet, the lead data is in CRM — and the two systems talk to each other after rest.
Calculate the cost of implementation
Select the scope of JSONMeet implementation in the company. The one-time price includes installation on your server, configuration of modules and team training. Your monthly subscription covers hosting, security patches, and updates.
Chcesz rozmowy z klientami na własnym serwerze, z AI po Twojej stronie?
Pokażemy Ci JSONMeet na żywo: wejdziesz jako gość przez poczekalnię, zobaczysz suflera od strony handlowca. 30 minut, bez prezentacji sprzedażowej.
Check before we talk
Services associated with this deployment. Take a look at the offer, see how we approach similar problems in other customers.
Podobał Ci się ten artykuł?
Jeśli po lekturze czujesz, że w Twojej firmie też są procesy warte przebudowy, umów bezpłatną rozmowę. Sprawdzimy razem, gdzie realnie wycieka sprzedaż i co ma sens wdrożyć w pierwszej kolejności.
Schedule a call