Compare commits

...
8 Commits
Author SHA1 Message Date
bj 9d33c2ee10 Pinned python package versions 2023-09-21 15:21:17 +03:00
bj aa17dcc48a Updated README.md to include ImageMagick 2023-09-21 15:14:38 +03:00
bj 6c1ece44ab Updated README.md with curl command and tesseract dependency 2023-09-21 15:02:06 +03:00
bj 3f3cc276f2 revised prompt to generate more cards 2023-09-21 14:52:06 +03:00
bj 34511cca74 Optimized prompt for brevity and clarity 2023-09-21 14:44:21 +03:00
bj 963d99404e BUGFIX: image processing only handles filenames with jpg 2023-09-21 14:42:47 +03:00
bj e96f23ddc8 revised README.md 2023-09-11 21:04:51 +03:00
bj 28e6c8d611 decoupled 2023-09-11 20:35:55 +03:00
10 changed files with 264 additions and 242 deletions
+92 -45
View File
@@ -1,67 +1,114 @@
# AnkiAI
# AnkiAI - Automated Anki Deck Creator
AnkiAI is a robust system that converts images containing text into structured Anki cards using Optical Character Recognition (OCR) and OpenAI's GPT-4 language model. Users can quickly generate decks of flashcards from their images for effective study.
AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's powerful natural language processing capabilities to automatically generate Anki decks from images containing text.
## Features
- Converts image content to textual content using OCR.
- Uses OpenAI's GPT-4 model to structure the content into Anki decks and cards.
- Outputs the structured content as an Anki package.
### Overview
## Dependencies
- genanki: Used for creating Anki decks and cards.
- Pillow: Image processing library.
- openai: API library for OpenAI's GPT-4 model.
- flask: Web server to host the service.
- AnkiAI is designed to streamline the process of creating Anki decks from images.
- The core idea is to use OCR to extract text from images and then use GPT-3 to transform this text into a structured Anki deck format.
- Users can make a POST request to a Flask server endpoint with their images to receive the Anki deck (.apkg file).
## Setup and Installation
### Directory Structure
- `.vscode/`: Contains configuration for VSCode debugger for Flask applications.
- `ankiai.py`: The main script that drives the creation of Anki decks from images.
- `constants.py`: Contains constant variables used across the project.
- `deck_creation.py`: Contains logic for communicating with OpenAI's API and deck creation using genanki.
- `image_processing.py`: Processes images, converting them for OCR and then performing OCR to extract text.
- `logging_config.py`: Logging configuration for the entire project.
- `server.py`: Flask server that provides an API endpoint to upload images and get back an Anki deck.
### Requirements
#### ImageMagick
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
```bash
sudo apt-get update
sudo apt-get install imagemagick
```
#### Tesseract
You need Tesseract for the OCR functionality:
```bash
sudo apt-get install tesseract-ocr
```
### Python Dependencies
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
You can install the Python dependencies via `pip` using the `requirements.txt` file:
```bash
pip install -r requirements.txt
```
### How to Run
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
1. Clone this repository:
```bash
git clone https://git.rudefox.io/bj/anki-json2ankicards.git
cd json2ankicards
export OPENAI_API_KEY=sk-myapikey
```
2. Set up a virtual environment and activate it:
```bash
python3 -m venv venv
source venv/bin/activate
```
2. **Run the Flask server**:
3. Install the required packages:
```bash
pip install -r requirements.txt
```
4. Set up the OpenAI API key:
```bash
export OPENAI_API_KEY=your_openai_api_key
```
5. Run the server:
```bash
python server.py
```
## Usage
This will start the Flask server. You can then make a POST request to `http://localhost:5000/deck-from-images` with your images to get an Anki deck.
1. Start the server as mentioned above.
3. **Run Directly**:
2. Use a tool like [Postman](https://www.postman.com/) or `curl` to send images to `http://localhost:5000/deck-from-images` as a multi-part POST request.
If you prefer not to use the Flask server, you can also run `ankiai.py` directly:
3. The server will respond with a downloadable Anki package. Import this into your Anki app and start studying!
```bash
python ankiai.py <directory_path_containing_images>
```
## Modules
### Example curl commands to interact with the service:
1. **ankiai.py**: The main module that orchestrates the flow.
2. **images2text.py**: Converts image content into text using OCR.
3. **json2deck.py**: Converts structured JSON data into an Anki package.
4. **prompt4cards.py**: Uses OpenAI to structure the content into Anki decks and cards.
5. **server.py**: Flask server to host the service.
You can make POST requests to the server using curl. Here are some examples from the command line history:
## Contributing
```bash
curl -X POST -o deck.apkg \
-F "image=@/home/ubuntu/Pictures/image1.png" \
-F "image=@/home/ubuntu/Pictures/image2.png" \
-F "image=@/home/ubuntu/Pictures/image3.png" \
http://localhost:5000/deck-from-images
```
Contributions are welcome! Please submit a pull request or open an issue to discuss changes or fixes.
Batch processing of images:
## License
```bash
for file in /home/ubuntu/Pictures/*; do
if [[ -f "$file" ]]; then
basefile=$(basename "$file");
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
fi;
done
```
[MIT License](LICENSE)
### How to Debug (VSCode Users)
- Open the project in VSCode.
- Set up your breakpoints.
- Use the VSCode debugger and select "Python: Flask" to start debugging the Flask server.
### Important Notes
- **API Key**: For the project to work, it is essential to have the `OPENAI_API_KEY` environment variable set.
- **Image Types**: Currently, the image processing module supports PNG, JPG, and JPEG formats.
- **Output**: The output `.apkg` file (Anki package file) will be named `out.apkg`.
### Acknowledgements
This project heavily relies on the `openai` library for processing and the `genanki` library for deck generation.
### Contributions
Contributions are always welcome. Please create a new issue or a pull request for any bug fixes or feature requests.
+9 -8
View File
@@ -2,18 +2,18 @@ import sys
import logging
from logging_config import setup_logging
from images2text import main as ocr_images
from prompt4cards import prompt_for_card_content, response_to_json
from json2deck import to_package
from image_processing import process_images
from deck_creation import prompt_for_card_content, response_to_json, to_package
APKG_FILE = "out.apkg"
setup_logging()
def images_to_package(directory_path, outfile):
ocr_text = ocr_images(directory_path)
def images_to_package(directory_path):
ocr_text = process_images(directory_path)
response_text = prompt_for_card_content(ocr_text)
deck_json = response_to_json(response_text)
to_package(deck_json).write_to_file(outfile)
logging.info(f"Deck created at: {outfile}")
return to_package(deck_json)
if __name__ == "__main__":
@@ -21,4 +21,5 @@ if __name__ == "__main__":
print("Usage: python ankiai.py <directory_path_containing_images>")
sys.exit(1)
images_to_package(sys.argv[1])
images_to_package(sys.argv[1]).write_to_file(APKG_FILE)
logging.info(f"Deck created at: {APKG_FILE}")
+4 -2
View File
@@ -1,8 +1,10 @@
# File and Directory Constants
IMAGE_KEY="image"
APKG_FILE="out.apkg"
CONVERTED_DIR = "converted"
FINAL_OUTPUT = "final.txt"
TEXT_OCR_FILE = "final.txt"
IMAGE_EXTENSIONS = ['.png', '.jpg', '.jpeg']
OUTPUT_FILENAME = "output_deck.json"
DECK_JSON_FILE = "output_deck.json"
# API Constants
API_KEY_ENV = "OPENAI_API_KEY"
+133
View File
@@ -0,0 +1,133 @@
import openai
import os
import json
import genanki
from logging_config import setup_logging
from constants import API_KEY_ENV, CHAT_MODEL
setup_logging()
API_KEY = os.environ.get(API_KEY_ENV)
if not API_KEY:
raise ValueError("Please set the OPENAI_API_KEY environment variable.")
openai.api_key = API_KEY
PROMPT_TEMPLATE = """
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
1. Every card should have a title, a question on the front, and an answer on the back.
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
6. Each index card should home in on answering a distinct question.
7. Limit each index card answer to no more than three sentences for brevity and clarity.
Structure your output as:
```
Deck Title: [Title of the Deck]
Cards:
- Title: [Card Title 1]
Front: [Question 1]
Back: [Answer 1]
- Title: [Card Title 2]
Front: [Question 2]
Back: [Answer 2]
... continue in this pattern
```
Content for reference:
{content}
"""
def prompt_for_card_content(text_content):
# Prepare the prompt
prompt = PROMPT_TEMPLATE.format(content=text_content)
# Get completion from the OpenAI ChatGPT API
response = openai.ChatCompletion.create(
model=CHAT_MODEL,
messages=[
{"role": "user", "content": prompt}
],
temperature=0,
)
# Extract content from response and save to a new file
return response.choices[0]['message']['content']
def response_to_json(response_text):
lines = [line.strip() for line in response_text.split("\n") if line.strip()]
deck_title = None
cards = []
current_card = {}
for line in lines:
if "Deck Title:" in line and not deck_title:
deck_title = line.split("Deck Title:", 1)[1].strip()
elif "Title:" in line:
if current_card: # If there's a card being processed, add it to cards
cards.append(current_card)
current_card = {}
current_card["Title"] = line.split("Title:", 1)[1].strip()
elif "Front:" in line:
current_card["Question"] = line.split("Front:", 1)[1].strip()
elif "Back:" in line:
current_card["Answer"] = line.split("Back:", 1)[1].strip()
if current_card: # Add the last card if it exists
cards.append(current_card)
return {
"DeckTitle": deck_title,
"Cards": cards
}
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607372319,
"Simple Model",
fields=[
{"name": "Title"},
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "{{Title}}",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def json_file_to_package(json_path):
with open(json_path, 'r', encoding='utf-8') as f:
json_data = json.load(f)
package = to_package(json_data)
return package
def to_package(deck_json):
deck_title = deck_json["DeckTitle"]
deck = genanki.Deck(1607372319, deck_title)
for card_json in deck_json["Cards"]:
title = card_json["Title"]
question = card_json["Question"]
answer = card_json["Answer"]
note = genanki.Note(
model=MY_MODEL,
fields=[title, question, answer]
)
deck.add_note(note)
return genanki.Package(deck)
+19 -7
View File
@@ -5,13 +5,21 @@ import logging
from logging_config import setup_logging
from subprocess import run, CalledProcessError
from concurrent.futures import ThreadPoolExecutor
from utilities import is_image_file, ensure_directory_exists
from constants import CONVERTED_DIR, FINAL_OUTPUT
from constants import CONVERTED_DIR, TEXT_OCR_FILE, IMAGE_EXTENSIONS
setup_logging()
def is_image_file(path):
return any(path.lower().endswith(ext) for ext in IMAGE_EXTENSIONS)
def ensure_directory_exists(directory):
if not os.path.exists(directory):
os.mkdir(directory)
def convert_image(image_path):
logging.info(f"Converting {image_path}...")
converted_path = os.path.join(CONVERTED_DIR, os.path.basename(image_path))
@@ -36,7 +44,11 @@ def convert_image(image_path):
def ocr_image(image_path):
logging.info(f"OCR'ing {image_path}...")
text_filename = os.path.basename(image_path).replace(".jpg", ".txt")
base_name = os.path.basename(image_path)
root_name, _ = os.path.splitext(base_name)
text_filename = f"{root_name}.txt"
text_path = os.path.join(CONVERTED_DIR, text_filename)
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
try:
@@ -62,7 +74,7 @@ def process_image(image_path):
return None
def main(directory_path):
def process_images(directory_path):
final_text = []
ensure_directory_exists(CONVERTED_DIR)
@@ -80,10 +92,10 @@ def main(directory_path):
# Filter out any None values and write the text to final.txt
final_text = [text for text in final_text if text is not None]
with open(FINAL_OUTPUT, 'w') as f:
with open(TEXT_OCR_FILE, 'w') as f:
f.write("\n".join(final_text))
logging.info(f"All images processed! Final output saved to {FINAL_OUTPUT}")
logging.info(f"All images processed! Final output saved to {TEXT_OCR_FILE}")
return final_text # Add this line
@@ -91,4 +103,4 @@ if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python images2text.py <directory_path>")
sys.exit(1)
main(sys.argv[1])
process_images(sys.argv[1])
-61
View File
@@ -1,61 +0,0 @@
import json
import genanki
import sys
import logging
from logging_config import setup_logging
setup_logging()
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607372319,
"Simple Model",
fields=[
{"name": "Title"},
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "{{Title}}",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def json_file_to_package(json_path):
with open(json_path, 'r', encoding='utf-8') as f:
json_data = json.load(f)
package = to_package(json_data)
return package
def to_package(deck_json):
deck_title = deck_json["DeckTitle"]
deck = genanki.Deck(1607372319, deck_title)
for card_json in deck_json["Cards"]:
title = card_json["Title"]
question = card_json["Question"]
answer = card_json["Answer"]
note = genanki.Note(
model=MY_MODEL,
fields=[title, question, answer]
)
deck.add_note(note)
return genanki.Package(deck)
if __name__ == "__main__":
if len(sys.argv) != 3:
print("Usage: python convert.py <input_json> <output_apkg>")
sys.exit(1)
input_json = sys.argv[1]
output_apkg = sys.argv[2]
json_file_to_package(input_json).write_to_file(output_apkg)
logging.info(f"Deck created at: {output_apkg}")
-103
View File
@@ -1,103 +0,0 @@
import openai
import os
import sys
import json
from constants import API_KEY_ENV, CHAT_MODEL, OUTPUT_FILENAME
API_KEY = os.environ.get(API_KEY_ENV)
if not API_KEY:
raise ValueError("Please set the OPENAI_API_KEY environment variable.")
openai.api_key = API_KEY
# Given prompt template
PROMPT_TEMPLATE = """
Please come up with a title for the deck and a set of 10 index cards for memorization,
including a title, front, and back for each card. The index cards should completely
capture the main points and themes of the text. In addition, they should contain any
numbers or data that humans might find difficult to remember. The goal of the index
card set is that one who memorizes it can provide a summary of the text to someone
else, conveying the main points and themes.
You will provide the deck title, and the titles, questions, and answers for each card
in a structured format as follows:
```
Deck Title: Title of the Deck
Cards:
- Title: Card Title 1
Front: What is the capital of New York?
Back: Albany
- Title: Card Title 2
Front: Where in the world is Carmen San Diego?
Back: Nobody knows
```
{content}
"""
def prompt_for_card_content(text_content):
# Prepare the prompt
prompt = PROMPT_TEMPLATE.format(content=text_content)
# Get completion from the OpenAI ChatGPT API
response = openai.ChatCompletion.create(
model=CHAT_MODEL,
messages=[
{"role": "user", "content": prompt}
],
temperature=0,
)
# Extract content from response and save to a new file
return response.choices[0]['message']['content']
def response_to_json(response_text):
lines = [line.strip() for line in response_text.split("\n") if line.strip()]
deck_title = None
cards = []
current_card = {}
for line in lines:
if "Deck Title:" in line and not deck_title:
deck_title = line.split("Deck Title:", 1)[1].strip()
elif "Title:" in line:
if current_card: # If there's a card being processed, add it to cards
cards.append(current_card)
current_card = {}
current_card["Title"] = line.split("Title:", 1)[1].strip()
elif "Front:" in line:
current_card["Question"] = line.split("Front:", 1)[1].strip()
elif "Back:" in line:
current_card["Answer"] = line.split("Back:", 1)[1].strip()
if current_card: # Add the last card if it exists
cards.append(current_card)
return {
"DeckTitle": deck_title,
"Cards": cards
}
if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python prompt4cards.py <text_file_path>")
sys.exit(1)
text_file_path = sys.argv[1]
# Read the text content
with open(text_file_path, 'r') as file:
text_content = file.read()
response_text = prompt_for_card_content(text_content)
deck_json = response_to_json(response_text)
with open(OUTPUT_FILENAME, 'w') as json_file:
json.dump(deck_json, json_file)
print(f"Saved generated deck to {OUTPUT_FILENAME}")
+3 -3
View File
@@ -1,4 +1,4 @@
genanki==0.8.0
Pillow
openai
flask
Pillow==10.0.1
openai==0.28.0
Flask==2.3.3
+7 -7
View File
@@ -3,17 +3,16 @@ import tempfile
import shutil
import logging
from logging_config import setup_logging
from flask import Flask, request, send_from_directory, jsonify
from werkzeug.utils import secure_filename
from ankiai import images_to_package
from constants import IMAGE_KEY, OUTPUT_FILE, NO_IMAGE_PART_ERROR, NO_SELECTED_FILE_ERROR, INVALID_FILENAME_ERROR
setup_logging()
from constants import IMAGE_KEY, APKG_FILE, NO_IMAGE_PART_ERROR, NO_SELECTED_FILE_ERROR, INVALID_FILENAME_ERROR
from logging_config import setup_logging
setup_logging()
app = Flask(__name__)
def save_uploaded_images(images, directory):
@@ -41,8 +40,9 @@ def deck_from_images():
save_uploaded_images(images, temp_dir)
try:
images_to_package(temp_dir, OUTPUT_FILE)
return send_from_directory('.', OUTPUT_FILE, as_attachment=True)
images_to_package(temp_dir).write_to_file(APKG_FILE)
logging.info(f"Anki package written to {APKG_FILE}")
return send_from_directory('.', APKG_FILE, as_attachment=True)
except Exception as e:
logging.error("Exception occurred: "+str(e), exc_info=True)
return jsonify({'error': str(e)}), 500
-9
View File
@@ -1,9 +0,0 @@
import os
from constants import IMAGE_EXTENSIONS
def is_image_file(path):
return any(path.lower().endswith(ext) for ext in IMAGE_EXTENSIONS)
def ensure_directory_exists(directory):
if not os.path.exists(directory):
os.mkdir(directory)