Prawbot. 3.9 million documents of Polish law in one conversation.
Prawbot. 3,9 mln dokumentów
polskiego prawa w jednej rozmowie.
- Ogólne chatboty zmyślają sygnatury i przepisy
- Dane klienta kancelarii lądują w obcej chmurze
- Brzmienie przepisu sprzed nowelizacji
- Praca na 50+ plikach naraz, nie na jednym pytaniu
- Własna baza prawa: 14 zbiorów, wyszukiwanie hybrydowe
- Czat ze sprawą i strażnicy cytowań
- Głęboki research, agenci, pamięć sprawy
- Model AI na własnych GPU, serwer w Warszawie
- 3,9 mln dokumentów w 14 bazach
- Każdy cytat sprawdzany u źródła przed odpowiedzią
- Płatność za użycie, bez abonamentu
- Od pierwszego commita do produkcji: 5 miesięcy
Prawnik pyta chatbota. Chatbot odpowiada pewnie. Sygnatura nie istnieje.
Kancelarie próbowały używać ogólnych asystentów AI do researchu prawnego. Efekt powtarzał się w kółko: odpowiedź brzmi jak z podręcznika, sygnatura wyroku wygląda poprawnie, a w bazie orzeczeń takiego wyroku nie ma. Albo jest, ale mówi o czymś innym. Prawnik musi sprawdzić wszystko od zera, więc oszczędność czasu znika.
Drugi problem: dane klienta. Umowa, akt oskarżenia czy dokumentacja sporu wklejona do publicznego chatbota trafia do infrastruktury, nad którą kancelaria nie ma kontroli. Dla zawodu objętego tajemnicą to nie jest szczegół techniczny, tylko pytanie o odpowiedzialność.
Trzeci problem: aktualność. Model językowy pamięta przepis w brzmieniu z dnia treningu. Polskie prawo zmienia się co tydzień. Bez dostępu do tekstu jednolitego „na dziś” i statusu aktu (obowiązuje, uchylony, zmieniony) odpowiedź może być poprawna dla stanu sprzed roku.
Czwarty problem jest najmniej widoczny w demo, a najważniejszy w pracy: kancelaria nie zadaje pojedynczych pytań. Dostaje sprawę z 50 plikami, skanami i historią korespondencji. Potrzebuje narzędzia, które to przeczyta, zapamięta i będzie na tym pracować przez tygodnie.
Mieliśmy przewagę na starcie. Wcześniej zbudowaliśmy dla własnych potrzeb bazę polskiego prawa z wyszukiwaniem semantycznym i zestawem narzędzi, z których korzystały nasze agenty AI. Postanowiliśmy zrobić z tego produkt dla kancelarii.
Co zbudowaliśmy. Cztery moduły Prawbota.
Baza polskiego prawa z wyszukiwaniem hybrydowym
14 zbiorów w jednym indeksie: orzeczenia sądów powszechnych, SN, TK i KIO, orzeczenia NSA i WSA, artykuły ustaw i kodeksów z historią brzmień, interpretacje podatkowe, decyzje UODO, KNF i UOKiK, KRS, akty prawne z metadanymi ELI, prace i głosowania Sejmu, zamówienia publiczne. Wyszukiwanie łączy dopasowanie semantyczne z pełnotekstowym i reranking wytrenowany na polskim języku prawniczym.
Czat ze sprawą i strażnicy cytowań
Model nie odpowiada z pamięci. Sięga do bazy przez ponad 40 narzędzi: szukaj wyroków, czytaj artykuł na wskazany dzień, sprawdź status aktu, pokaż powiązane orzeczenia. Strażnicy cytowań pilnują, żeby każda sygnatura i każdy artykuł w odpowiedzi były realnie przeczytane w tej rozmowie. Cytat, którego nie da się zakotwiczyć w źródle, jest oznaczany jako wymagający weryfikacji.
Głęboki research, agenci i pamięć sprawy
Dla trudnych spraw: planowanie pytań badawczych, równoległe agenty-specjaliści od poszczególnych dziedzin, weryfikacja kompletności i synteza w długą opinię z przypisami. Pamięć sprawy trzyma ustalenia między rozmowami. Konstruktor przepływów agentowych pozwala kancelarii złożyć własny proces na dziesiątkach plików. Generator pism z ponad 400 typami dokumentów domyka pracę.
Platforma SaaS dla kancelarii
Konta osób prywatnych i kancelarii, weryfikacja kancelarii po NIP i numerze wpisu, role w workspace, sprawy i pliki z automatycznym odczytem skanów, rozliczenie za faktyczne użycie z kredytami na start, płatności online. Komplet dokumentów: regulamin, Privacy Policy, umowa powierzenia (art. 28 RODO), informacja o AI zgodna z art. 50 AI Act, procedura realizacji praw osób.
Pod maską
Cała architektura opiera się na wzorcu RAG z weryfikacją cytatów. Backend w Pythonie (FastAPI), frontend w React i TypeScript. Indeks prawa w bazie wektorowej Qdrant: 14 kolekcji, ponad 20 mln fragmentów tekstu, każdy z wektorem semantycznym i indeksem pełnotekstowym. Narzędzia dla modelu wystawione przez Model Context Protocol (MCP), więc ten sam zestaw obsługuje czat w aplikacji, agentów researchu i zewnętrzne klienty MCP.
Model językowy to model o otwartych wagach, uruchomiony na naszych własnych serwerach GPU. Żadne zapytanie ani dokument nie trafia do zewnętrznego dostawcy modeli. Platforma i dane stoją na dedykowanym serwerze w centrum danych OVHcloud w Warszawie. Wdrożenia idą w trybie blue/green za reverse proxy Caddy, pipeline CI/CD na Gitea Actions uruchamia testy i przełącza ruch bez przestoju.
Ogólny chatbot a Prawbot. Ta sama sprawa, inna odpowiedź.
- Sygnatura wyroku z pamięci modelu, bez sprawdzenia
- Przepis w brzmieniu z dnia treningu
- Dokumenty klienta w infrastrukturze poza kontrolą kancelarii
- Jedno pytanie, jedna odpowiedź, zero pamięci sprawy
- Abonament per użytkownik, niezależnie od użycia
- Brak umowy powierzenia dopasowanej do kancelarii
- Sygnatura przeczytana w bazie w tej rozmowie, albo oznaczona do weryfikacji
- Przepis na wskazany dzień, ze statusem aktu i nowelizacjami
- Dane na serwerze w Warszawie, model na własnych GPU
- Sprawa z plikami, pamięcią i historią researchu
- Płatność za faktyczne użycie, kredyty na start
- Jedna umowa RODO: administrator i dostawca w jednym
Liczby projektu.
w tym 744 tys. orzeczeń
semantyczny + pełnotekstowy
~200 tys. linii kodu
kod i izolacja klientów
Do tego: 100+ migracji bazy danych, 120+ plików testów, ponad 400 typów pism w generatorze, 10 dokumentów prawnych gotowych do publikacji. Wszystko w jednym repozytorium, z wiki architektury pisaną równolegle z kodem.
Najtrudniejsze nie było postawienie modelu. Najtrudniejsze było nauczyć system, żeby nie odpowiadał, dopóki nie przeczyta źródła. Prawnik nie potrzebuje asystenta, który brzmi pewnie. Potrzebuje takiego, który pokazuje, skąd to wziął.
Jak wyglądała budowa. Pięć miesięcy.
1
Baza prawa, narzędzia dla modelu, weryfikator cytatów
Indeksowanie 14 zbiorów do Qdrant, wyszukiwanie hybrydowe z rerankingiem po polsku, warstwa narzędzi w protokole MCP. Pierwszy moduł SaaS: weryfikator, który wyciąga z odpowiedzi sygnatury i artykuły i sprawdza je w bazie. 270 commitów w pierwszym miesiącu.
2
Aplikacja dla kancelarii
Konta, role, workspace, sprawy i pliki. Czat ze sprawą ze strumieniowaniem odpowiedzi i wywołaniami narzędzi. Odczyt skanów i zdjęć dokumentów. Pierwsza wersja rozliczeń za tokeny.
3
Deep research, agents, audits
Research pipeline: question plan, specialist agents, completeness verification, synthesis. Memory of the case. Agent flow builder and template library. Two security audits (code and client isolation), fixes implemented. Landing, price list, set of legal documents.
4
Production and own equipment
Dedicated server in Warsaw, blue/green deploy, CI/CD. Transferring the language model to your own GPU servers with context counted in hundreds of thousands of tokens. Reasoning modes selected separately for chat and research.
5
Benchmarks, tests and gatekeepers
We built own set of legal benchmarks: cases from several fields, with traps for specific provisions and fictitious theses of judgments, assessed automatically by an independent model judge and compared with the best commercial models. The conclusion was clear: errors appeared almost exclusively where the model responded without reaching the database. We included citation and regulation guards, sifted out amending laws from the results, and banned guessing at facts in the prompt. Onboarding of subsequent law firms with verification within 24 hours.
AI for a law firm: how to implement an assistant that doesn't make up case numbers
A guide for law firm partners, in-house counsel, and people responsible for IT at legal businesses. We explain how AI for a law firm based on RAG and Polish law works, how it differs from a general-purpose chatbot, and what to ask a vendor before implementation.
Why a general-purpose chatbot isn't suitable for legal research
A general-purpose AI assistant answers from the model’s memory. It sounds confident and gives a case number in the correct format, but the case-law database has no such ruling, or it concerns something else entirely. The lawyer then has to check everything from scratch, so the time savings disappear.
The second problem is currency. The model remembers a provision as worded on the date of training, and Polish law changes constantly. Without access to the consolidated text as of a given date and the act's status, an answer may be correct for the state of the law a year ago.
The third problem is organizational. A law firm rarely asks one-off questions. It works on a case with dozens of files, scans, and correspondence it keeps coming back to for weeks. A tool built for one-off questions can't handle that.
RAG on Polish law: how it works and what it requires
RAG (retrieval-augmented generation) is a pattern in which the model searches for and reads sources before answering, and only then writes the answer based on them. In law, this means the quality of the system depends primarily on the database and the search engine, and only secondarily on the model itself.
In the implementation described above, the database covers 3.9 million documents in 14 collections, including 744,000 court rulings. These include rulings of the common courts, the Supreme Court, the Constitutional Tribunal, the National Appeals Chamber (KIO), and the administrative courts (NSA and WSA), articles of statutes and codes with their wording history, tax rulings, decisions of the UODO, KNF, and UOKiK regulators, legal acts with ELI metadata, and the work of the Sejm. The index contains more than 20 million text fragments.
Collecting the data alone isn't enough. Legal language requires hybrid search: semantic matching captures the meaning of the question, while full-text search catches exact case numbers and institution names. Finally, a reranker trained on Polish legal language orders the results by relevance. The database also has to be refreshed regularly from official sources, in this case daily.
The practical takeaway: when asking a vendor about RAG on Polish law, ask first about the scope and freshness of the database and how search works. The name of the language model matters less, because the model can be swapped, while building a well-indexed database of case law and legislation is the most labor-intensive part of the project.
AI hallucinations in law: verifying citations at the source
RAG reduces the risk of hallucinations but doesn't eliminate it. The model can still add a case number it hasn't read. That's why Prawbot has an additional layer: citation guards compare every case number and every article in the answer against the list of documents actually read in that conversation. A citation that can't be anchored in a source is flagged as requiring verification.
This mechanism came out of testing. The team built its own set of legal benchmarks with case problems from several areas of law, including traps involving special provisions and fabricated case-law holdings. The conclusion was clear: errors occurred almost exclusively where the model answered without consulting the database.
A practical tip when choosing a tool: ask to see what happens when the system can't find a source. Good AI for a law firm should clearly flag that, not answer in a confident tone.
Attorney-client privilege and a language model on your own servers
A contract, an indictment, or litigation documents pasted into a public chatbot end up on infrastructure the law firm has no control over. For a profession bound by confidentiality, that's a question of liability, not a technical detail.
The solution is an open-weight model running on servers controlled by the service provider or the law firm itself. In Prawbot, the model runs on its own GPU servers, and the platform and data sit on a dedicated server in an OVHcloud data center in Warsaw. Queries and documents are not sent to an external model provider. For law firms with stricter requirements, a dedicated instance isolated from the shared cloud is available.
On top of that comes the formal layer. During implementation, it's worth checking whether the vendor has a data processing agreement compliant with Art. 28 GDPR, an AI disclosure compliant with Art. 50 of the AI Act, and a procedure for handling data subject rights. In the project described, a full set of such documents was part of the product, and the code and client isolation were verified by two security audits.
What AI for a law firm should be able to do beyond answering questions
Chat with citations is just the beginning. In day-to-day work, what counts are features that support the entire case:
- a provision as of a given date, with the status of the act (in force, repealed, amended) and a list of amendments,
- case memory, which retains findings between conversations,
- reading scans and photos of documents uploaded to the case,
- deep research: a plan of research questions, parallel specialist agents for individual areas of law, and synthesis into a memo with footnotes,
- legal document generator, in Prawbot over 400 document types,
- custom agent workflows, in which the law firm assembles a repeatable process across dozens of files.
Technically, all of these features use a single set of more than 40 tools exposed via the Model Context Protocol (MCP): search judgments, read an article as of a given date, check the status of an act, show related rulings. As a result, the chat, research agents, and external MCP clients all work on the same verified sources.
How to choose an AI implementation partner for a law firm
Before signing a contract, it’s worth asking the vendor a few specific questions:
- Does the system read the source before answering, and how does it flag citations it hasn't verified?
- What data was it tested on, and can you test it on our documents?
- Where do the model and database physically run, and who is the data processor?
- How often is the legal database refreshed, and where does the data come from?
- What does maintenance look like after implementation, not just the project itself?
The answers to these questions tell you more than a demo. A demo shows well-chosen examples, but in practice, a tool's usefulness depends on how it behaves on a difficult case with incomplete data.
We describe the ready-made platform on the page Prawbot: an AI platform for lawyers, and we bring the same RAG pattern with source verification to other companies' documentation and procedures as part of AI implementations, and the easiest way to start a conversation about your specific law firm is through the contact.
Three questions we hear from every law firm.
AI makes things up. How do I know that this judgment really exists and says what the assistant claims?
Prawbot does not quote from memory. Before providing the reference number, he reads the ruling in the database in the same conversation. The citation guard compares each reference number and each article in the response with the list of actually read documents. If something can't be anchored, you get a clear „requires review” tag instead of a confident tone. Each quote is clickable and leads to the full text. More about the mechanism in our article about the six signals of a hallucinating chatbot.
Where is my customer data and who has access to it?
The platform and data are located on a dedicated server in the OVHcloud data center in Warsaw (ISO/IEC 27001 and 27701). The language model runs on our own GPU servers, so the content of conversations and documents does not go to any external AI provider. Prawbot is the sole data controller, OVHcloud acts as a processor only on our instructions. You sign the entrustment agreement once, with us. We don't train models on your questions or documents.
How much does it cost if there is no subscription?
You pay for actual use, measured in input and output tokens. Fragments of the conversation that the model already has in its cache are not recounted. The law firm receives start-up loans after verification, and a private person immediately after registration. You can see consumption on an ongoing basis in the panel, per user and per case. We separately implement Prawbot as a platform in the law firm - you can calculate the approximate range using the calculator below.
Does Prawbot work in the on-premise version, on our own server?
Yes. We install Prawbot Enterprise as an instance dedicated to the law firm, on your server or on a dedicated server of a selected cloud provider. The language model, Qdrant vector database and backend run in an isolated environment, without connection to the JSON Crew cloud. This is a scenario for large law firms and corporate legal departments with strict compliance requirements.
Where does Prawbot get its current regulations? Does it take into account the latest amendments?
The vector database is refreshed every day from official sources: the Sejm, the Journal of Laws, systems for publishing the jurisprudence of common courts, the Supreme Court, the Constitutional Tribunal, the Supreme Administrative Court and the Provincial Administrative Court, tax interpretations and decisions of the Personal Data Protection Office (UODO/KNF/UOKiK). Each article of the act has a full wording history with ELI metadata - Prawbot can show the provision as of „as of” and indicate that there were X amendments after that date.
How long does it take to implement Prawbot in a law firm?
Standard implementation (access to the general law database + integration with the office system + team training) takes 2-3 weeks from order to production. Implementation with the law firm's own database of precedents means an additional 2-4 weeks for indexing and verification of search quality. We set a specific date after a 30-minute diagnostic conversation and counting the volume of documents.
How to choose an AI implementer so as not to end up with a system that no one in the company uses?
Three signals that we check in-house and that are worth asking every supplier about: (1) own benchmarks on your documents, not on generic sets, (2) architecture with source citation verification, not trust in the model, (3) post-implementation maintenance plan, not just the design. We write about this in more detail in the article „how to choose an AI implementer for the legal industry”.
Calculate the cost of implementation
Choose the scope of Prawbot implementation in your law firm. The one-time price includes knowledge base setup, integration and training. The monthly subscription covers access for the entire law firm and updates to the Polish law database.
Masz bazę wiedzy, na której AI ma odpowiadać bez zmyślania?
Pokażemy Ci na żywo, jak Prawbot sprawdza cytaty u źródła, i jak ten sam wzorzec przenosimy na dokumentację, procedury albo katalog produktów w Twojej firmie. 30 minut, bez prezentacji sprzedażowej.
Check before we talk
Services associated with this deployment. Take a look at the offer, see how we approach similar problems in other customers.
Podobał Ci się ten artykuł?
Jeśli po lekturze czujesz, że w Twojej firmie też są procesy warte przebudowy, umów bezpłatną rozmowę. Sprawdzimy razem, gdzie realnie wycieka sprzedaż i co ma sens wdrożyć w pierwszej kolejności.
Schedule a call