Compare commits

..
8 Commits
Author SHA1 Message Date
bj 9d33c2ee10 Pinned python package versions 2023-09-21 15:21:17 +03:00
bj aa17dcc48a Updated README.md to include ImageMagick 2023-09-21 15:14:38 +03:00
bj 6c1ece44ab Updated README.md with curl command and tesseract dependency 2023-09-21 15:02:06 +03:00
bj 3f3cc276f2 revised prompt to generate more cards 2023-09-21 14:52:06 +03:00
bj 34511cca74 Optimized prompt for brevity and clarity 2023-09-21 14:44:21 +03:00
bj 963d99404e BUGFIX: image processing only handles filenames with jpg 2023-09-21 14:42:47 +03:00
bj e96f23ddc8 revised README.md 2023-09-11 21:04:51 +03:00
bj 28e6c8d611 decoupled 2023-09-11 20:35:55 +03:00
10 changed files with 264 additions and 242 deletions
+83 -36
View File
@@ -1,67 +1,114 @@
# AnkiAI # AnkiAI - Automated Anki Deck Creator
AnkiAI is a robust system that converts images containing text into structured Anki cards using Optical Character Recognition (OCR) and OpenAI's GPT-4 language model. Users can quickly generate decks of flashcards from their images for effective study. AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's powerful natural language processing capabilities to automatically generate Anki decks from images containing text.
## Features ### Overview
- Converts image content to textual content using OCR.
- Uses OpenAI's GPT-4 model to structure the content into Anki decks and cards.
- Outputs the structured content as an Anki package.
## Dependencies - AnkiAI is designed to streamline the process of creating Anki decks from images.
- genanki: Used for creating Anki decks and cards. - The core idea is to use OCR to extract text from images and then use GPT-3 to transform this text into a structured Anki deck format.
- Pillow: Image processing library. - Users can make a POST request to a Flask server endpoint with their images to receive the Anki deck (.apkg file).
- openai: API library for OpenAI's GPT-4 model.
- flask: Web server to host the service.
## Setup and Installation ### Directory Structure
- `.vscode/`: Contains configuration for VSCode debugger for Flask applications.
- `ankiai.py`: The main script that drives the creation of Anki decks from images.
- `constants.py`: Contains constant variables used across the project.
- `deck_creation.py`: Contains logic for communicating with OpenAI's API and deck creation using genanki.
- `image_processing.py`: Processes images, converting them for OCR and then performing OCR to extract text.
- `logging_config.py`: Logging configuration for the entire project.
- `server.py`: Flask server that provides an API endpoint to upload images and get back an Anki deck.
### Requirements
#### ImageMagick
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
1. Clone this repository:
```bash ```bash
git clone https://git.rudefox.io/bj/anki-json2ankicards.git sudo apt-get update
cd json2ankicards sudo apt-get install imagemagick
``` ```
2. Set up a virtual environment and activate it: #### Tesseract
You need Tesseract for the OCR functionality:
```bash ```bash
python3 -m venv venv sudo apt-get install tesseract-ocr
source venv/bin/activate
``` ```
### Python Dependencies
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
You can install the Python dependencies via `pip` using the `requirements.txt` file:
3. Install the required packages:
```bash ```bash
pip install -r requirements.txt pip install -r requirements.txt
``` ```
4. Set up the OpenAI API key: ### How to Run
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
```bash ```bash
export OPENAI_API_KEY=your_openai_api_key export OPENAI_API_KEY=sk-myapikey
``` ```
5. Run the server: 2. **Run the Flask server**:
```bash ```bash
python server.py python server.py
``` ```
## Usage This will start the Flask server. You can then make a POST request to `http://localhost:5000/deck-from-images` with your images to get an Anki deck.
1. Start the server as mentioned above. 3. **Run Directly**:
2. Use a tool like [Postman](https://www.postman.com/) or `curl` to send images to `http://localhost:5000/deck-from-images` as a multi-part POST request. If you prefer not to use the Flask server, you can also run `ankiai.py` directly:
3. The server will respond with a downloadable Anki package. Import this into your Anki app and start studying! ```bash
python ankiai.py <directory_path_containing_images>
```
## Modules ### Example curl commands to interact with the service:
1. **ankiai.py**: The main module that orchestrates the flow. You can make POST requests to the server using curl. Here are some examples from the command line history:
2. **images2text.py**: Converts image content into text using OCR.
3. **json2deck.py**: Converts structured JSON data into an Anki package.
4. **prompt4cards.py**: Uses OpenAI to structure the content into Anki decks and cards.
5. **server.py**: Flask server to host the service.
## Contributing ```bash
curl -X POST -o deck.apkg \
-F "image=@/home/ubuntu/Pictures/image1.png" \
-F "image=@/home/ubuntu/Pictures/image2.png" \
-F "image=@/home/ubuntu/Pictures/image3.png" \
http://localhost:5000/deck-from-images
```
Contributions are welcome! Please submit a pull request or open an issue to discuss changes or fixes. Batch processing of images:
## License ```bash
for file in /home/ubuntu/Pictures/*; do
if [[ -f "$file" ]]; then
basefile=$(basename "$file");
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
fi;
done
```
[MIT License](LICENSE) ### How to Debug (VSCode Users)
- Open the project in VSCode.
- Set up your breakpoints.
- Use the VSCode debugger and select "Python: Flask" to start debugging the Flask server.
### Important Notes
- **API Key**: For the project to work, it is essential to have the `OPENAI_API_KEY` environment variable set.
- **Image Types**: Currently, the image processing module supports PNG, JPG, and JPEG formats.
- **Output**: The output `.apkg` file (Anki package file) will be named `out.apkg`.
### Acknowledgements
This project heavily relies on the `openai` library for processing and the `genanki` library for deck generation.
### Contributions
Contributions are always welcome. Please create a new issue or a pull request for any bug fixes or feature requests.
+9 -8
View File
@@ -2,18 +2,18 @@ import sys
import logging import logging
from logging_config import setup_logging from logging_config import setup_logging
from images2text import main as ocr_images from image_processing import process_images
from prompt4cards import prompt_for_card_content, response_to_json from deck_creation import prompt_for_card_content, response_to_json, to_package
from json2deck import to_package
APKG_FILE = "out.apkg"
setup_logging() setup_logging()
def images_to_package(directory_path, outfile): def images_to_package(directory_path):
ocr_text = ocr_images(directory_path) ocr_text = process_images(directory_path)
response_text = prompt_for_card_content(ocr_text) response_text = prompt_for_card_content(ocr_text)
deck_json = response_to_json(response_text) deck_json = response_to_json(response_text)
to_package(deck_json).write_to_file(outfile) return to_package(deck_json)
logging.info(f"Deck created at: {outfile}")
if __name__ == "__main__": if __name__ == "__main__":
@@ -21,4 +21,5 @@ if __name__ == "__main__":
print("Usage: python ankiai.py <directory_path_containing_images>") print("Usage: python ankiai.py <directory_path_containing_images>")
sys.exit(1) sys.exit(1)
images_to_package(sys.argv[1]) images_to_package(sys.argv[1]).write_to_file(APKG_FILE)
logging.info(f"Deck created at: {APKG_FILE}")
+4 -2
View File
@@ -1,8 +1,10 @@
# File and Directory Constants # File and Directory Constants
IMAGE_KEY="image"
APKG_FILE="out.apkg"
CONVERTED_DIR = "converted" CONVERTED_DIR = "converted"
FINAL_OUTPUT = "final.txt" TEXT_OCR_FILE = "final.txt"
IMAGE_EXTENSIONS = ['.png', '.jpg', '.jpeg'] IMAGE_EXTENSIONS = ['.png', '.jpg', '.jpeg']
OUTPUT_FILENAME = "output_deck.json" DECK_JSON_FILE = "output_deck.json"
# API Constants # API Constants
API_KEY_ENV = "OPENAI_API_KEY" API_KEY_ENV = "OPENAI_API_KEY"
+133
View File
@@ -0,0 +1,133 @@
import openai
import os
import json
import genanki
from logging_config import setup_logging
from constants import API_KEY_ENV, CHAT_MODEL
setup_logging()
API_KEY = os.environ.get(API_KEY_ENV)
if not API_KEY:
raise ValueError("Please set the OPENAI_API_KEY environment variable.")
openai.api_key = API_KEY
PROMPT_TEMPLATE = """
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
1. Every card should have a title, a question on the front, and an answer on the back.
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
6. Each index card should home in on answering a distinct question.
7. Limit each index card answer to no more than three sentences for brevity and clarity.
Structure your output as:
```
Deck Title: [Title of the Deck]
Cards:
- Title: [Card Title 1]
Front: [Question 1]
Back: [Answer 1]
- Title: [Card Title 2]
Front: [Question 2]
Back: [Answer 2]
... continue in this pattern
```
Content for reference:
{content}
"""
def prompt_for_card_content(text_content):
# Prepare the prompt
prompt = PROMPT_TEMPLATE.format(content=text_content)
# Get completion from the OpenAI ChatGPT API
response = openai.ChatCompletion.create(
model=CHAT_MODEL,
messages=[
{"role": "user", "content": prompt}
],
temperature=0,
)
# Extract content from response and save to a new file
return response.choices[0]['message']['content']
def response_to_json(response_text):
lines = [line.strip() for line in response_text.split("\n") if line.strip()]
deck_title = None
cards = []
current_card = {}
for line in lines:
if "Deck Title:" in line and not deck_title:
deck_title = line.split("Deck Title:", 1)[1].strip()
elif "Title:" in line:
if current_card: # If there's a card being processed, add it to cards
cards.append(current_card)
current_card = {}
current_card["Title"] = line.split("Title:", 1)[1].strip()
elif "Front:" in line:
current_card["Question"] = line.split("Front:", 1)[1].strip()
elif "Back:" in line:
current_card["Answer"] = line.split("Back:", 1)[1].strip()
if current_card: # Add the last card if it exists
cards.append(current_card)
return {
"DeckTitle": deck_title,
"Cards": cards
}
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607372319,
"Simple Model",
fields=[
{"name": "Title"},
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "{{Title}}",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def json_file_to_package(json_path):
with open(json_path, 'r', encoding='utf-8') as f:
json_data = json.load(f)
package = to_package(json_data)
return package
def to_package(deck_json):
deck_title = deck_json["DeckTitle"]
deck = genanki.Deck(1607372319, deck_title)
for card_json in deck_json["Cards"]:
title = card_json["Title"]
question = card_json["Question"]
answer = card_json["Answer"]
note = genanki.Note(
model=MY_MODEL,
fields=[title, question, answer]
)
deck.add_note(note)
return genanki.Package(deck)
+19 -7
View File
@@ -5,13 +5,21 @@ import logging
from logging_config import setup_logging from logging_config import setup_logging
from subprocess import run, CalledProcessError from subprocess import run, CalledProcessError
from concurrent.futures import ThreadPoolExecutor from concurrent.futures import ThreadPoolExecutor
from utilities import is_image_file, ensure_directory_exists from constants import CONVERTED_DIR, TEXT_OCR_FILE, IMAGE_EXTENSIONS
from constants import CONVERTED_DIR, FINAL_OUTPUT
setup_logging() setup_logging()
def is_image_file(path):
return any(path.lower().endswith(ext) for ext in IMAGE_EXTENSIONS)
def ensure_directory_exists(directory):
if not os.path.exists(directory):
os.mkdir(directory)
def convert_image(image_path): def convert_image(image_path):
logging.info(f"Converting {image_path}...") logging.info(f"Converting {image_path}...")
converted_path = os.path.join(CONVERTED_DIR, os.path.basename(image_path)) converted_path = os.path.join(CONVERTED_DIR, os.path.basename(image_path))
@@ -36,7 +44,11 @@ def convert_image(image_path):
def ocr_image(image_path): def ocr_image(image_path):
logging.info(f"OCR'ing {image_path}...") logging.info(f"OCR'ing {image_path}...")
text_filename = os.path.basename(image_path).replace(".jpg", ".txt")
base_name = os.path.basename(image_path)
root_name, _ = os.path.splitext(base_name)
text_filename = f"{root_name}.txt"
text_path = os.path.join(CONVERTED_DIR, text_filename) text_path = os.path.join(CONVERTED_DIR, text_filename)
cmd = ["tesseract", image_path, text_path.replace(".txt", "")] cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
try: try:
@@ -62,7 +74,7 @@ def process_image(image_path):
return None return None
def main(directory_path): def process_images(directory_path):
final_text = [] final_text = []
ensure_directory_exists(CONVERTED_DIR) ensure_directory_exists(CONVERTED_DIR)
@@ -80,10 +92,10 @@ def main(directory_path):
# Filter out any None values and write the text to final.txt # Filter out any None values and write the text to final.txt
final_text = [text for text in final_text if text is not None] final_text = [text for text in final_text if text is not None]
with open(FINAL_OUTPUT, 'w') as f: with open(TEXT_OCR_FILE, 'w') as f:
f.write("\n".join(final_text)) f.write("\n".join(final_text))
logging.info(f"All images processed! Final output saved to {FINAL_OUTPUT}") logging.info(f"All images processed! Final output saved to {TEXT_OCR_FILE}")
return final_text # Add this line return final_text # Add this line
@@ -91,4 +103,4 @@ if __name__ == "__main__":
if len(sys.argv) != 2: if len(sys.argv) != 2:
print("Usage: python images2text.py <directory_path>") print("Usage: python images2text.py <directory_path>")
sys.exit(1) sys.exit(1)
main(sys.argv[1]) process_images(sys.argv[1])
-61
View File
@@ -1,61 +0,0 @@
import json
import genanki
import sys
import logging
from logging_config import setup_logging
setup_logging()
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607372319,
"Simple Model",
fields=[
{"name": "Title"},
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "{{Title}}",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def json_file_to_package(json_path):
with open(json_path, 'r', encoding='utf-8') as f:
json_data = json.load(f)
package = to_package(json_data)
return package
def to_package(deck_json):
deck_title = deck_json["DeckTitle"]
deck = genanki.Deck(1607372319, deck_title)
for card_json in deck_json["Cards"]:
title = card_json["Title"]
question = card_json["Question"]
answer = card_json["Answer"]
note = genanki.Note(
model=MY_MODEL,
fields=[title, question, answer]
)
deck.add_note(note)
return genanki.Package(deck)
if __name__ == "__main__":
if len(sys.argv) != 3:
print("Usage: python convert.py <input_json> <output_apkg>")
sys.exit(1)
input_json = sys.argv[1]
output_apkg = sys.argv[2]
json_file_to_package(input_json).write_to_file(output_apkg)
logging.info(f"Deck created at: {output_apkg}")
-103
View File
@@ -1,103 +0,0 @@
import openai
import os
import sys
import json
from constants import API_KEY_ENV, CHAT_MODEL, OUTPUT_FILENAME
API_KEY = os.environ.get(API_KEY_ENV)
if not API_KEY:
raise ValueError("Please set the OPENAI_API_KEY environment variable.")
openai.api_key = API_KEY
# Given prompt template
PROMPT_TEMPLATE = """
Please come up with a title for the deck and a set of 10 index cards for memorization,
including a title, front, and back for each card. The index cards should completely
capture the main points and themes of the text. In addition, they should contain any
numbers or data that humans might find difficult to remember. The goal of the index
card set is that one who memorizes it can provide a summary of the text to someone
else, conveying the main points and themes.
You will provide the deck title, and the titles, questions, and answers for each card
in a structured format as follows:
```
Deck Title: Title of the Deck
Cards:
- Title: Card Title 1
Front: What is the capital of New York?
Back: Albany
- Title: Card Title 2
Front: Where in the world is Carmen San Diego?
Back: Nobody knows
```
{content}
"""
def prompt_for_card_content(text_content):
# Prepare the prompt
prompt = PROMPT_TEMPLATE.format(content=text_content)
# Get completion from the OpenAI ChatGPT API
response = openai.ChatCompletion.create(
model=CHAT_MODEL,
messages=[
{"role": "user", "content": prompt}
],
temperature=0,
)
# Extract content from response and save to a new file
return response.choices[0]['message']['content']
def response_to_json(response_text):
lines = [line.strip() for line in response_text.split("\n") if line.strip()]
deck_title = None
cards = []
current_card = {}
for line in lines:
if "Deck Title:" in line and not deck_title:
deck_title = line.split("Deck Title:", 1)[1].strip()
elif "Title:" in line:
if current_card: # If there's a card being processed, add it to cards
cards.append(current_card)
current_card = {}
current_card["Title"] = line.split("Title:", 1)[1].strip()
elif "Front:" in line:
current_card["Question"] = line.split("Front:", 1)[1].strip()
elif "Back:" in line:
current_card["Answer"] = line.split("Back:", 1)[1].strip()
if current_card: # Add the last card if it exists
cards.append(current_card)
return {
"DeckTitle": deck_title,
"Cards": cards
}
if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python prompt4cards.py <text_file_path>")
sys.exit(1)
text_file_path = sys.argv[1]
# Read the text content
with open(text_file_path, 'r') as file:
text_content = file.read()
response_text = prompt_for_card_content(text_content)
deck_json = response_to_json(response_text)
with open(OUTPUT_FILENAME, 'w') as json_file:
json.dump(deck_json, json_file)
print(f"Saved generated deck to {OUTPUT_FILENAME}")
+3 -3
View File
@@ -1,4 +1,4 @@
genanki==0.8.0 genanki==0.8.0
Pillow Pillow==10.0.1
openai openai==0.28.0
flask Flask==2.3.3
+7 -7
View File
@@ -3,17 +3,16 @@ import tempfile
import shutil import shutil
import logging import logging
from logging_config import setup_logging
from flask import Flask, request, send_from_directory, jsonify from flask import Flask, request, send_from_directory, jsonify
from werkzeug.utils import secure_filename from werkzeug.utils import secure_filename
from ankiai import images_to_package from ankiai import images_to_package
from constants import IMAGE_KEY, OUTPUT_FILE, NO_IMAGE_PART_ERROR, NO_SELECTED_FILE_ERROR, INVALID_FILENAME_ERROR from constants import IMAGE_KEY, APKG_FILE, NO_IMAGE_PART_ERROR, NO_SELECTED_FILE_ERROR, INVALID_FILENAME_ERROR
setup_logging()
from logging_config import setup_logging from logging_config import setup_logging
setup_logging()
app = Flask(__name__) app = Flask(__name__)
def save_uploaded_images(images, directory): def save_uploaded_images(images, directory):
@@ -41,8 +40,9 @@ def deck_from_images():
save_uploaded_images(images, temp_dir) save_uploaded_images(images, temp_dir)
try: try:
images_to_package(temp_dir, OUTPUT_FILE) images_to_package(temp_dir).write_to_file(APKG_FILE)
return send_from_directory('.', OUTPUT_FILE, as_attachment=True) logging.info(f"Anki package written to {APKG_FILE}")
return send_from_directory('.', APKG_FILE, as_attachment=True)
except Exception as e: except Exception as e:
logging.error("Exception occurred: "+str(e), exc_info=True) logging.error("Exception occurred: "+str(e), exc_info=True)
return jsonify({'error': str(e)}), 500 return jsonify({'error': str(e)}), 500
-9
View File
@@ -1,9 +0,0 @@
import os
from constants import IMAGE_EXTENSIONS
def is_image_file(path):
return any(path.lower().endswith(ext) for ext in IMAGE_EXTENSIONS)
def ensure_directory_exists(directory):
if not os.path.exists(directory):
os.mkdir(directory)