Compare commits

...
6 Commits
4 changed files with 73 additions and 28 deletions
+46 -8
View File
@@ -20,16 +20,27 @@ AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's
### Requirements
To run AnkiAI, you'll need to have the following dependencies installed:
#### ImageMagick
```
genanki==0.8.0
Pillow
openai
flask
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
```bash
sudo apt-get update
sudo apt-get install imagemagick
```
You can install these via `pip` using the `requirements.txt` file:
#### Tesseract
You need Tesseract for the OCR functionality:
```bash
sudo apt-get install tesseract-ocr
```
### Python Dependencies
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
You can install the Python dependencies via `pip` using the `requirements.txt` file:
```bash
pip install -r requirements.txt
@@ -38,7 +49,11 @@ pip install -r requirements.txt
### How to Run
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
```bash
export OPENAI_API_KEY=sk-myapikey
```
2. **Run the Flask server**:
```bash
@@ -55,6 +70,29 @@ pip install -r requirements.txt
python ankiai.py <directory_path_containing_images>
```
### Example curl commands to interact with the service:
You can make POST requests to the server using curl. Here are some examples from the command line history:
```bash
curl -X POST -o deck.apkg \
-F "image=@/home/ubuntu/Pictures/image1.png" \
-F "image=@/home/ubuntu/Pictures/image2.png" \
-F "image=@/home/ubuntu/Pictures/image3.png" \
http://localhost:5000/deck-from-images
```
Batch processing of images:
```bash
for file in /home/ubuntu/Pictures/*; do
if [[ -f "$file" ]]; then
basefile=$(basename "$file");
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
fi;
done
```
### How to Debug (VSCode Users)
- Open the project in VSCode.
+19 -16
View File
@@ -15,28 +15,31 @@ if not API_KEY:
openai.api_key = API_KEY
# Given prompt template
PROMPT_TEMPLATE = """
Please come up with a title for the deck and a set of 10 index cards for memorization,
including a title, front, and back for each card. The index cards should completely
capture the main points and themes of the text. In addition, they should contain any
numbers or data that humans might find difficult to remember. The goal of the index
card set is that one who memorizes it can provide a summary of the text to someone
else, conveying the main points and themes.
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
You will provide the deck title, and the titles, questions, and answers for each card
in a structured format as follows:
1. Every card should have a title, a question on the front, and an answer on the back.
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
6. Each index card should home in on answering a distinct question.
7. Limit each index card answer to no more than three sentences for brevity and clarity.
Structure your output as:
```
Deck Title: Title of the Deck
Deck Title: [Title of the Deck]
Cards:
- Title: Card Title 1
Front: What is the capital of New York?
Back: Albany
- Title: Card Title 2
Front: Where in the world is Carmen San Diego?
Back: Nobody knows
- Title: [Card Title 1]
Front: [Question 1]
Back: [Answer 1]
- Title: [Card Title 2]
Front: [Question 2]
Back: [Answer 2]
... continue in this pattern
```
Content for reference:
{content}
"""
+5 -1
View File
@@ -44,7 +44,11 @@ def convert_image(image_path):
def ocr_image(image_path):
logging.info(f"OCR'ing {image_path}...")
text_filename = os.path.basename(image_path).replace(".jpg", ".txt")
base_name = os.path.basename(image_path)
root_name, _ = os.path.splitext(base_name)
text_filename = f"{root_name}.txt"
text_path = os.path.join(CONVERTED_DIR, text_filename)
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
try:
+3 -3
View File
@@ -1,4 +1,4 @@
genanki==0.8.0
Pillow
openai
flask
Pillow==10.0.1
openai==0.28.0
Flask==2.3.3