Can AI replace Harvey?
You can absolutely build a local RAG assistant over your own PDFs in an afternoon, and for reading your lease or a vendor contract that is genuinely enough. Harvey is not sold on that loop. It is sold on trained-and-evaluated legal workflows, curated case law and regulatory sources under license, deployment that survives a law firm's security review, and the ability to put a name behind an output that a partner will bill against. The thing you cannot one-shot is the confidence to rely on the answer, which in legal work is the entire product. A personal replacement is fine for personal stakes and dangerous the moment money or a filing depends on it.
01What it costs
Checked Aug 18, 2026 · source: harvey.ai.
02Could AI build it for you?
The core job: Indexes a folder of your own contracts and PDFs locally, then answers questions about them with quotes and page citations so you can check every claim yourself.
What a working version needs:
- An LLM API key, or a local model via Ollama if you would rather nothing leaves the machine
- Your own documents: no licensed case law, no statute databases, no primary sources
- Python 3.11 and a willingness to read the cited passage rather than trust the summary
The retrieval half is a weekend, the reliance half is not. The moat is licensed primary law plus an enterprise security posture that took years to clear, and a name a partner can put in a risk memo.
03What you'd give up
- Licensed primary law: case law, statutes, filings and regulatory corpora you cannot legally scrape together
- Workflow products that have been evaluated by actual lawyers: diligence checklists, redline review, deposition prep
- Firm-grade deployment: SSO, data residency, retention controls, audit logs, security questionnaires answered
- Anyone to blame. Your hallucination is your malpractice exposure
- Integration into the systems legal work actually lives in: DMS, iManage, Word, the review platform
A law firm is not paying for text generation, it is paying for defensibility. The output has to be traceable to a licensed source, the deployment has to pass a security review that takes months, the workflows have to have been tested against how associates actually do diligence, and there has to be a vendor contract with indemnities when something goes wrong. Individual lawyers also cannot use a homemade tool on client matters without answering awkward questions about where the privileged data went. A local RAG box over your own files solves none of that and does not need to; it solves reading your own documents faster, which is a different and much smaller job.
05The build prompt
Paste this into an AI coding tool (such as Claude, ChatGPT, Lovable or Replit) to build your own version. Read the verdict first: this one is hard to get right.
Build a local document Q&A tool for my own contracts and PDFs. Python 3.11, single project, no web framework, no accounts, no telemetry. Stack, no alternatives: - CLI with Typer - pypdf for text extraction, page numbers preserved - SQLite with the sqlite-vec extension for vector storage - OpenAI API for embeddings and answers, key from .env via python-dotenv, plus an --ollama flag that swaps to a local model at http://localhost:11434 Commands: - ingest PATH: walk a folder, extract text per page, chunk to roughly 800 tokens with 100 overlap, store chunk text plus doc name plus page number, embed and index. Skip files already ingested unless --force. - ask "QUESTION": retrieve top 12 chunks, then answer with an LLM that is instructed to answer only from the provided chunks and to say "not in these documents" when the answer is absent. Every claim must carry an inline citation like [contract.pdf p.4]. - sources "QUESTION": print the retrieved chunks verbatim with file and page, no LLM, so I can read the raw text. - docs: list ingested files, page counts, chunk counts. In scope: local-only storage in ./index.db, deterministic chunking, a --model flag, plain text output. Out of scope: web UI, multi-user, cloud sync, any bundled case law or statute data, any attempt to cite external legal sources. Print a one-line disclaimer after every answer: this is a reading aid over my own files, not legal advice, verify each citation. Include a README with setup, .env.example, and a short section explaining that answers are only as good as the documents I ingested.
App prices, verdicts, alternatives and build prompts are adapted from Can I Vibecode It? (MIT License, © 2026 Rob Hallam). Each price shows the date it was checked and its source. Prices change; confirm on the vendor's site before you decide.
Get new verdicts in your inbox.
One short email when new verdicts land: what AI can now do for you, and what it still gets wrong. No spam. Unsubscribe anytime.