Can AI replace Cluely?
The core loop is genuinely thin: capture a screenshot, transcribe the mic, feed both to a multimodal model, stream the answer into a transparent always-on-top window on a hotkey. An agent gets that working in a session, and it will feel uncomfortably close to the demo. The gaps are the unglamorous parts: capturing the other side's audio takes virtual audio devices, latency has to beat the conversation, and the headline trick of staying invisible in screen shares depends on fragile platform window flags that vary by OS and conferencing app. Your DIY version will work fine when you are alone with your own screen, and may quietly betray you the one time it matters.
01What it costs
Checked Aug 16, 2026 · source: cluely.com.
02Could AI build it for you?
The core job: A local Electron overlay bound to a hotkey that screenshots your active display, transcribes recent audio, sends both to a multimodal LLM, and streams a short answer into a translucent window nobody else is supposed to see.
What a working version needs:
- An OpenAI API key (or any multimodal chat endpoint) in .env
- macOS or Windows desktop, plus screen recording and microphone permissions
- A virtual audio device for other-party audio: BlackHole on macOS, VB-Cable or WASAPI loopback on Windows
- Node 20 and willingness to sign or self-trust an unsigned Electron build
The value proposition is helping you fake competence, so budget for the possibility that your build works perfectly and you still get caught.
03What you'd give up
- Screen-share invisibility that has actually been tested against Zoom, Meet, Teams and OS screenshot tools
- Sub-second latency and the prompt tuning that makes answers short enough to read while someone is talking
- System audio capture that just works without you installing and routing a virtual audio device
- Mobile and browser-extension coverage, plus a hosted account so your setup follows you to another machine
- Somebody else's legal and PR exposure for a product whose pitch is 'cheat on everything'
Because the hard part is not the LLM call, it is the twenty small platform details that make an overlay silent, invisible and fast under pressure, and because people who reach for this tool are, by definition, not in the mood to debug a virtual audio driver ten minutes before an interview. Paying converts a fragile personal hack into something that mostly behaves on a laptop you did not configure yourself.
05The build prompt
Paste this into an AI coding tool (such as Claude, ChatGPT, Lovable or Replit) to build your own version.
Build a local desktop AI overlay assistant called "Peek". Stack: Electron 30 with TypeScript, Vite for the renderer, no framework beyond plain TS and CSS. Target macOS first, keep Windows paths behind a small platform module. Everything runs locally except LLM calls. Behavior: 1. On launch, create a frameless, transparent, always-on-top BrowserWindow, 420x520, positioned top-right of the primary display. No dock icon, no menu bar item beyond a tray icon with Quit. 2. Set the window to ignore mouse events by default; hold a modifier hotkey to make it interactive. 3. Global hotkeys: Cmd+Shift+Space asks a question from the current context, Cmd+Shift+H toggles visibility, Cmd+Shift+Enter follows up in the same thread. 4. On ask: capture a PNG of the active display via Electron desktopCapturer at max 1600px wide, grab the last 30 seconds of rolling audio transcript, and send both to the model. 5. Audio: capture from a configurable input device using the renderer's MediaRecorder in 5 second chunks, transcribe each chunk with OpenAI whisper-1, keep a rolling 60 second transcript buffer in memory only. Never write audio to disk. 6. LLM: call the OpenAI chat completions endpoint with a multimodal message (screenshot plus transcript plus user intent), stream tokens into the overlay. System prompt: answer in under 60 words, lead with the answer, bullets only when listing, no preamble. 7. Attempt screen-share exclusion: call setContentProtection(true) on the window, and on Windows use the SetWindowDisplayAffinity equivalent flag. Print a startup warning in the console that exclusion is best effort and must be verified manually. 8. Settings: a small JSON config at userData/config.json for model name, input device id, hotkeys, and transcript window length. No settings UI beyond a tray menu item that opens the file. In scope: the overlay, hotkeys, screenshot capture, rolling transcription, streaming answers, tray quit, README with macOS permission steps and BlackHole routing instructions for capturing other-party audio. Out of scope: accounts, login, cloud sync, telemetry, analytics, auto-update, code signing, mobile, browser extension, meeting integrations, transcript persistence, any database. Secrets: OPENAI_API_KEY in .env, loaded in the main process only. Never expose the key to the renderer; proxy all API calls through IPC. Commit a .env.example. Deliver: working npm scripts dev and build, a README with a 60 second setup, and a HONESTY.md file listing what will break, starting with screen-share invisibility and audio device routing.
App prices, verdicts, alternatives and build prompts are adapted from Can I Vibecode It? (MIT License, © 2026 Rob Hallam). Each price shows the date it was checked and its source. Prices change; confirm on the vendor's site before you decide.
Get new verdicts in your inbox.
One short email when new verdicts land: what AI can now do for you, and what it still gets wrong. No spam. Unsubscribe anytime.