Passive OCR and other 'AI' tools on the Linux desktop

Artemis_Mystique@lemmy.ml · 5 months ago

Passive OCR and other 'AI' tools on the Linux desktop

qaz@lemmy.world · edit-2 5 months ago

I personally have a script that allows me to press a keybinding and then select some part of the display after which it copies the contents to my clipboard as text. Are you looking for something like that?

Here it is, it should work on both Wayland and X11, but does require having spectacle and tesseract installed.

#!/bin/bash

# Written by u/qaz licensed under GPL3

if ! which spectacle > /dev/null; then
    kdialog --sorry "spectacle, the required screenshotting tool, is not installed."
    exit 1
fi

if ! which tesseract > /dev/null; then
    kdialog --sorry "tesseract, the required OCR package, is not installed."
    exit 1
fi

screenshot_tempfile=$(mktemp)
spectacle -brn -o $screenshot_tempfile
text_tempfile=$(mktemp)
tesseract $screenshot_tempfile $text_tempfile
rm $screenshot_tempfile

result_text=$(cat $text_tempfile.txt)
rm $text_tempfile.txt

# Copy to either X11 or Wayland clipboard
echo $result_text | xclip -selection clipboard
wl-copy "$result_text"

notify-send -u low -t 2500 "Copied text to clipboard" "$result_text"