Firmowa baza wiedzy z lokalną AI
Zadajesz pytanie. Dostajesz odpowiedź z Twoich dokumentów — ze źródłem.
RAG od local-ai.company łączy czat z wiedzą zapisaną w dokumentach firmy. Pliki zostają na komputerze w biurze, a każdy pracownik ma też swoją prywatną część, do której nikt inny nie zagląda.
Droga jednego pytania
- Pytaniepełnym zdaniem, w oknie czatu
- Uprawnieniasystem bierze pod uwagę tylko to, co ta osoba może widzieć
- Wyszukiwaniepo znaczeniu i po dokładnych słowach: numerach, datach, kwotach
- Rerankerdrugi model czyta pytanie razem z każdym fragmentem i ustala kolejność
- Odpowiedźz nazwą pliku — albo z uczciwym „nie wiem”
Company knowledge base with local AI
You ask a question. You get an answer from your own documents — with the source.
RAG by local-ai.company connects chat with the knowledge in your company's documents. The files stay on the computer in your office, and every employee also has a private part that nobody else can look into.
The path of one question
- Questiona full sentence in the chat window
- Permissionsonly what this person is allowed to see is considered
- Searchby meaning and by exact terms: numbers, dates, amounts
- Rerankera second model reads the question with each excerpt and sets the order
- Answerwith the file name — or an honest “I don't know”
Co potrafi
Wiedza firmy dostępna w rozmowie
Nie kolejny folder do przeszukiwania. Czat, który czyta Twoje dokumenty i odpowiada ich słowami.
Odpowiedzi ze źródłem
Każda odpowiedź wskazuje plik, z którego pochodzi. Łatwo sprawdzić, zanim się coś podpisze.
Uczciwe „nie wiem”
Gdy w dokumentach nie ma odpowiedzi, system mówi to wprost, zamiast składać zdanie z przypadkowego pliku.
Prywatne dane pracownika
Każdy ma własny folder i własną część bazy wiedzy — niewidoczną dla innych, także dla administratora.
Pamięć rozmów na własność
Ważne ustalenia z czatu trafiają do prywatnej bazy pracownika. Zostają jego, nawet po zmianie narzędzia.
Poczta i skany
Wiadomości e-mail, dokumenty Word i Excel, PDF-y oraz skany z rozpoznawaniem polskiego tekstu.
Pliki prosto w rozmowie
Dokument dołączony do czatu jest czytany od razu — bez czekania na nocną aktualizację.
Porządek sam się utrzymuje
Zmieniony plik zastępuje starą wersję, usunięty przestaje być cytowany. Bez ręcznego sprzątania.
Dokumenty w kilku językach
Pytanie po polsku przeszukuje też dokumenty angielskie. Gdy trafienia są słabe, asystent sam ponawia wyszukiwanie po angielsku.
Asystent bez dostępu do prywatnych danych
Agent AI, który dba o bazę wiedzy, korzysta z danych firmowych, ale nie widzi prywatnych folderów pracowników.
What it does
Company knowledge, available in a conversation
Not another folder to dig through. A chat that reads your documents and answers in their words.
Answers with sources
Every answer names the file it came from, so it is easy to check before anything gets signed.
An honest “I don't know”
When the documents don't contain the answer, the system says so instead of stitching a sentence together from an unrelated file.
Private employee data
Everyone has their own folder and their own part of the knowledge base — invisible to others, including the administrator.
Conversation memory they own
Key points from chats go into the employee's private knowledge base. They stay theirs, even after switching tools.
E-mail and scans
E-mail messages, Word and Excel files, PDFs and scans, with recognition of Polish text.
Files right in the chat
A document attached to a conversation is read immediately — no waiting for the overnight update.
It keeps itself tidy
A changed file replaces its old version, a deleted one stops being quoted. No manual clean-up.
Documents in several languages
A question in Polish also searches English documents. When the matches are weak, the assistant retries the search in English on its own.
An assistant without access to private data
The AI agent that looks after the knowledge base uses company data, but cannot see employees' private folders.
Reranker
Dlaczego samo wyszukiwanie nie wystarcza
Wyszukiwanie wektorowe porównuje pytanie z fragmentami dokumentów po znaczeniu. Jest szybkie, ale łatwo je zmylić: krótka notatka, w której pada szukane nazwisko, wygląda na trafniejszą niż długi dokument, w którym to samo nazwisko pojawia się raz. Pierwsze trafienia bywają „podobne, ale nie o tym”.
Reranker to drugi model. Czyta pytanie razem z każdym znalezionym fragmentem i dopiero wtedy ocenia, czy fragment naprawdę na nie odpowiada. Jest wolniejszy od wyszukiwarki, dlatego ocenia tylko kandydatów z pierwszego etapu — i ustawia ich w kolejności rzeczywistej trafności, zanim cokolwiek trafi do modelu rozmowy.
Szerokie sito
Wyszukiwanie po znaczeniu i po dokładnych słowach — numerach, datach, kwotach, kodach — we wszystkich zbiorach, do których pytający ma dostęp. Każdy zbiór ma zagwarantowane miejsce wśród kandydatów, więc kilkanaście kart pacjentów nie przegrywa z tysiącem krótkich rekordów.
Dokładne czytanie
Reranker ocenia każdego kandydata razem z pytaniem i układa kolejność. Do modelu rozmowy trafiają tylko najlepsze fragmenty, uzupełnione o sąsiedni tekst, żeby zdanie nie urywało się w połowie.
Uczciwy wynik
Słabe trafienia są oznaczane jako niepewne. Gdy wszystkie są słabe, asystent mówi wprost, że nie ma tej informacji, zamiast składać odpowiedź z przypadkowego pliku.
- Właściwy dokument w pierwszej piątce wyników
- 76% → 90% po kalibracji wyszukiwania i rerankera · 21 pytań kontrolnych · sierpień 2026
- Średnia pozycja właściwego dokumentu (MRR)
- 0,69 → 0,83 1,0 oznaczałoby, że właściwy dokument jest zawsze na pierwszym miejscu
- Wyszukiwanie z rerankingiem
- ~1,3 s karta graficzna 12 GB · wrzesień 2026
- Ten sam reranker bez karty graficznej
- 11–54 s starszy procesor serwerowy, zależnie od ustawień · dlatego do pracy na bieżąco zalecam kartę
Reranker to otwarty, wielojęzyczny model bge-reranker-v2-m3 (BAAI) na licencji Apache 2.0. Jego karta jest publiczna — możesz ją sprawdzić sam.
Reranker
Why search alone is not enough
Vector search compares a question with document excerpts by meaning. It is fast, but easy to mislead: a short note that mentions the surname you are looking for looks more relevant than a long document in which the same surname appears once. The top hits are often “similar, but not about that”.
A reranker is a second model. It reads the question together with each retrieved excerpt, and only then judges whether the excerpt really answers it. It is slower than the search engine, so it only scores the candidates from the first stage — and puts them in order of real relevance before anything reaches the chat model.
A wide net
Search by meaning and by exact terms — numbers, dates, amounts, codes — across every collection the person asking can access. Each collection has a guaranteed share of the candidates, so a dozen patient records don't lose out to a thousand short entries.
Close reading
The reranker scores each candidate together with the question and sets the order. Only the best excerpts reach the chat model, with the neighbouring text added so sentences aren't cut off halfway.
An honest result
Weak matches are flagged as uncertain. When all of them are weak, the assistant says plainly that it doesn't have the information, instead of assembling an answer from an unrelated file.
- Right document in the top five results
- 76% → 90% after calibrating search and the reranker · 21 control questions · August 2026
- Average rank of the right document (MRR)
- 0.69 → 0.83 1.0 would mean the right document always comes first
- Search with reranking
- ~1.3 s 12 GB graphics card · September 2026
- The same reranker without a graphics card
- 11–54 s older server CPU, depending on settings · which is why I recommend a card for day-to-day work
The reranker is the open, multilingual model bge-reranker-v2-m3 (BAAI), licensed under Apache 2.0. Its model card is public, so you can check it yourself.
Jak to wygląda w pracy
Trzy kroki, żadnej konfiguracji
Wrzuć plik do folderu
Zwykły folder sieciowy w Windows. Wspólny dla firmy albo prywatny — nazwany adresem e-mail pracownika.
Poczekaj do rana
W nocy system czyta nowe i zmienione dokumenty. Pliki dołączone do rozmowy działają od razu.
Zapytaj w czacie
Pełnym zdaniem, po ludzku. Odpowiedź przychodzi z cytatem i nazwą pliku.
How it works day to day
Three steps, no configuration
Drop a file into a folder
An ordinary network folder in Windows. Shared with the company, or private — named after the employee's e-mail address.
Wait until morning
Overnight, the system reads new and changed documents. Files attached to a chat work straight away.
Ask in the chat
In a full, natural sentence. The answer comes with a quote and the file name.
Skomplikowane w środku, proste na zewnątrz
Dlaczego to trudne do zbudowania — i dlaczego łatwe w obsłudze
Co musi działać pod spodem
Dziesiątki elementów, które muszą zgadzać się ze sobą
- Czytanie każdego rodzaju dokumentuTabele, pieczątki, skany słabej jakości, polskie znaki — każdy format wymaga innej ścieżki.
- Wyszukiwanie, które nie gubi małych zbiorówKilkanaście kart pacjentów nie może przegrać z tysiącem krótkich rekordów tylko dlatego, że jest ich mniej.
- Ocena trafności drugim modelemPierwsze wyszukiwanie znajduje kandydatów, reranker czyta je razem z pytaniem i układa kolejność.
- Kontrola dostępu w jednym miejscuKażde zapytanie przechodzi przez tę samą bramkę — i jest to sprawdzane testami, a nie obietnicą.
- Karta graficzna dzielona w czasieModel rozmowy, wyszukiwanie i nocna aktualizacja walczą o tę samą pamięć. System sam rozpoznaje kartę i dobiera tryb pracy.
- Bezpieczne usuwanieSkasowany dokument musi zniknąć z odpowiedzi — ale odłączony dysk sieciowy nie może wyglądać jak skasowanie wszystkiego.
- Pomiar zamiast wrażeniaKażda zmiana jest sprawdzana na pytaniach kontrolnych i testach izolacji danych, zanim trafi do klienta.
Co widzi pracownik
Folder i okno czatu
- Pliki zapisuje się tak jak zawsze — w folderze w Windows.
- Pyta się pełnym zdaniem, bez słów kluczowych i bez składni.
- Odpowiedź ma źródło, więc łatwo ją sprawdzić.
- Prywatny folder jest prywatny — bez ustawień i uprawnień do klikania.
- Własną część bazy można wyłączyć jednym plikiem w swoim folderze.
- Nic nie trzeba instalować na komputerach pracowników — wystarczy przeglądarka.
Complex inside, simple outside
Why it is hard to build — and why it is easy to use
What has to work underneath
Dozens of parts that must agree with each other
- Reading every kind of documentTables, stamps, poor-quality scans, Polish characters — each format needs its own path.
- Search that doesn't lose small collectionsA dozen patient records must not lose to a thousand short entries just because there are fewer of them.
- A second model that judges relevanceThe first search finds candidates; the reranker reads them together with the question and puts them in order.
- Access control in a single placeEvery query passes through the same gate — and that is proven by tests, not promised.
- A graphics card shared over timeThe chat model, search and the overnight update compete for the same memory. The system recognises the card and chooses how to work.
- Safe deletionA deleted document must disappear from answers — but a disconnected network drive must never look like everything was deleted.
- Measurement, not impressionsEvery change is checked against control questions and data-isolation tests before it reaches a client.
What an employee sees
A folder and a chat window
- Files are saved as always — in a Windows folder.
- Questions are asked in full sentences, with no keywords or syntax.
- Answers come with a source, so they are easy to verify.
- A private folder is private — no settings or permissions to click through.
- Your own part of the knowledge base can be switched off with a single file in your folder.
- Nothing to install on employees' computers — a web browser is enough.
Sprzęt
Działa z kartą, którą masz — albo bez niej
Dokumenty u Ciebie, model rozmowy u mnie
Czytanie dokumentów, indeks i wyszukiwanie działają na komputerze w Twoim biurze. Model rozmowy pracuje na moim węźle w Kielcach (4× RTX 3090), połączony szyfrowanym tunelem. Do modelu trafiają wyłącznie bieżące pytania z fragmentami potrzebnymi do odpowiedzi — i nie są zapisywane.
Bez karty wszystko działa, ale wyszukiwanie z rerankerem jest wtedy wyraźnie wolniejsze. Karta 12 GB wystarcza, bo zajmuje się tylko wyszukiwaniem.
Wszystko w firmie
Model rozmowy pracuje lokalnie w godzinach pracy (6–22), a w nocy (22–6) karta aktualizuje bazę wiedzy. Nic nie opuszcza biura.
Dwie karty albo osobny komputer na model pozwalają obu częściom pracować jednocześnie.
System sam rozpoznaje zainstalowaną kartę przy uruchomieniu i dobiera tryb pracy — wymiana karty nie wymaga zmian w ustawieniach. Im szybciej przejdziesz na własną kartę, tym szybciej żadne zapytanie nie wychodzi poza firmę. Ile daje która karta, pokazuje benchmark.
Hardware
Works with the graphics card you have — or without one
Documents with you, the chat model with me
Reading documents, the index and search all run on the computer in your office. The chat model runs on my node in Kielce (4× RTX 3090), over an encrypted tunnel. Only current questions and the excerpts needed to answer them reach the model — and they are not stored.
Everything works without a card, but search with the reranker is noticeably slower. A 12 GB card is enough, because it only handles search.
Everything in-house
The chat model runs locally during working hours (6:00–22:00), and at night (22:00–6:00) the card updates the knowledge base. Nothing leaves the office.
Two cards, or a separate machine for the model, let both parts work at the same time.
The system recognises the installed card at start-up and chooses its working mode — changing the card requires no change in settings. The sooner you move to your own card, the sooner no query leaves your company at all. The benchmark shows what each card delivers.
Sprawdźmy, czy taki RAG ma sens u Ciebie
Na konsultacji ustalimy, jakie dokumenty i procesy mają trafić do bazy wiedzy, kto ma widzieć które zbiory i jaki sprzęt wystarczy przy wielkości Twojej firmy.
Let's see whether this RAG makes sense for you
In a consultation we'll work out which documents and processes belong in the knowledge base, who should see which collections, and what hardware is enough for a company your size.