Compare commits

..
20 Commits
Author SHA1 Message Date
bj 9d33c2ee10 Pinned python package versions 2023-09-21 15:21:17 +03:00
bj aa17dcc48a Updated README.md to include ImageMagick 2023-09-21 15:14:38 +03:00
bj 6c1ece44ab Updated README.md with curl command and tesseract dependency 2023-09-21 15:02:06 +03:00
bj 3f3cc276f2 revised prompt to generate more cards 2023-09-21 14:52:06 +03:00
bj 34511cca74 Optimized prompt for brevity and clarity 2023-09-21 14:44:21 +03:00
bj 963d99404e BUGFIX: image processing only handles filenames with jpg 2023-09-21 14:42:47 +03:00
bj e96f23ddc8 revised README.md 2023-09-11 21:04:51 +03:00
bj 28e6c8d611 decoupled 2023-09-11 20:35:55 +03:00
bj a24d0ea96f extracted contants and added logging 2023-09-11 20:02:17 +03:00
bj c3bdf2dd5e overhauled the project to get away from files (a little) 2023-09-11 19:10:42 +03:00
bj a13f92548c refactorings and security enhancements 2023-09-08 18:22:57 +03:00
bj 61827b4e11 updated README.md 2023-09-08 16:29:15 +03:00
bj c3ca249877 Created a server to serve apkg files in respose to image posts 2023-09-08 16:24:34 +03:00
bj af4b615aa3 The pipeline works end-to-end 2023-09-08 12:58:19 +03:00
bj 91350f3ca3 Created a pipeline 2023-09-08 12:37:24 +03:00
bj 8213eff41b Externalized API Key 2023-09-08 12:32:49 +03:00
bj b4be376ec8 Updated README.md 2023-09-08 12:24:37 +03:00
bj ff06c093e3 Can convert from text to csv list of questions 2023-09-08 12:19:33 +03:00
bj 51401ba964 Updated README.md 2023-09-08 11:30:05 +03:00
bj 18a9cb0dd9 Can convert, OCR and combine image files to text in parallel 2023-09-08 11:25:59 +03:00
11 changed files with 469 additions and 87 deletions
+1
View File
@@ -2,3 +2,4 @@ venv/
*.pyc
__pycache__/
.env
+27
View File
@@ -0,0 +1,27 @@
{
// Use IntelliSense to learn about possible attributes.
// Hover to view descriptions of existing attributes.
// For more information, visit: https://go.microsoft.com/fwlink/?linkid=830387
"version": "0.2.0",
"configurations": [
{
"name": "Python: Flask",
"type": "python",
"request": "launch",
"module": "flask",
"envFile": "${workspaceFolder}/.env",
"env": {
"FLASK_APP": "server.py",
"FLASK_ENV": "development",
"FLASK_DEBUG": "0"
},
"args": [
"run",
"--no-debugger",
"--no-reload"
],
"jinja": true,
"justMyCode": true
}
]
}
+93 -38
View File
@@ -1,59 +1,114 @@
# csv2ankicards
# AnkiAI - Automated Anki Deck Creator
A simple tool to convert CSV files into Anki deck packages (.apkg files).
AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's powerful natural language processing capabilities to automatically generate Anki decks from images containing text.
## Features
### Overview
- Converts a CSV file with questions and answers into an Anki deck package.
- There are only two columns in the CSV file, separated by the first comma encountered.
- CSV files should have a "Front" column for questions and a "Back" column for answers.
- AnkiAI is designed to streamline the process of creating Anki decks from images.
- The core idea is to use OCR to extract text from images and then use GPT-3 to transform this text into a structured Anki deck format.
- Users can make a POST request to a Flask server endpoint with their images to receive the Anki deck (.apkg file).
## Installation
### Directory Structure
1. Clone this repository:
```bash
git clone https://git.rudefox.io/bj/anki-csv2ankicards.git
cd csv2ankicards
```
- `.vscode/`: Contains configuration for VSCode debugger for Flask applications.
- `ankiai.py`: The main script that drives the creation of Anki decks from images.
- `constants.py`: Contains constant variables used across the project.
- `deck_creation.py`: Contains logic for communicating with OpenAI's API and deck creation using genanki.
- `image_processing.py`: Processes images, converting them for OCR and then performing OCR to extract text.
- `logging_config.py`: Logging configuration for the entire project.
- `server.py`: Flask server that provides an API endpoint to upload images and get back an Anki deck.
2. Set up a virtual environment and activate it:
```bash
python3 -m venv venv
source venv/bin/activate
```
### Requirements
3. Install the required packages:
```bash
pip install -r requirements.txt
```
#### ImageMagick
## Usage
To convert a CSV file into an Anki deck package:
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
```bash
python csv2ankicards.py /path/to/your/csvfile.csv output.apkg
sudo apt-get update
sudo apt-get install imagemagick
```
This will produce an `output.apkg` file which can then be imported into Anki.
#### Tesseract
### CSV Format
The CSV file should follow this format:
You need Tesseract for the OCR functionality:
```bash
sudo apt-get install tesseract-ocr
```
Front,Back
Your question here,Your answer here, and here
Another question,list of: answer1, answer2, answer3
...
### Python Dependencies
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
You can install the Python dependencies via `pip` using the `requirements.txt` file:
```bash
pip install -r requirements.txt
```
**Note:** If your answers contain commas, they will be considered as part of the answer. Only the first comma is used to separate the question from the answer.
### How to Run
## License
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
[MIT License](LICENSE)
```bash
export OPENAI_API_KEY=sk-myapikey
```
## Contributing
2. **Run the Flask server**:
Pull requests are welcome. For major changes, please open an issue first to discuss what you would like to change.
```bash
python server.py
```
This will start the Flask server. You can then make a POST request to `http://localhost:5000/deck-from-images` with your images to get an Anki deck.
3. **Run Directly**:
If you prefer not to use the Flask server, you can also run `ankiai.py` directly:
```bash
python ankiai.py <directory_path_containing_images>
```
### Example curl commands to interact with the service:
You can make POST requests to the server using curl. Here are some examples from the command line history:
```bash
curl -X POST -o deck.apkg \
-F "image=@/home/ubuntu/Pictures/image1.png" \
-F "image=@/home/ubuntu/Pictures/image2.png" \
-F "image=@/home/ubuntu/Pictures/image3.png" \
http://localhost:5000/deck-from-images
```
Batch processing of images:
```bash
for file in /home/ubuntu/Pictures/*; do
if [[ -f "$file" ]]; then
basefile=$(basename "$file");
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
fi;
done
```
### How to Debug (VSCode Users)
- Open the project in VSCode.
- Set up your breakpoints.
- Use the VSCode debugger and select "Python: Flask" to start debugging the Flask server.
### Important Notes
- **API Key**: For the project to work, it is essential to have the `OPENAI_API_KEY` environment variable set.
- **Image Types**: Currently, the image processing module supports PNG, JPG, and JPEG formats.
- **Output**: The output `.apkg` file (Anki package file) will be named `out.apkg`.
### Acknowledgements
This project heavily relies on the `openai` library for processing and the `genanki` library for deck generation.
### Contributions
Contributions are always welcome. Please create a new issue or a pull request for any bug fixes or feature requests.
+25
View File
@@ -0,0 +1,25 @@
import sys
import logging
from logging_config import setup_logging
from image_processing import process_images
from deck_creation import prompt_for_card_content, response_to_json, to_package
APKG_FILE = "out.apkg"
setup_logging()
def images_to_package(directory_path):
ocr_text = process_images(directory_path)
response_text = prompt_for_card_content(ocr_text)
deck_json = response_to_json(response_text)
return to_package(deck_json)
if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python ankiai.py <directory_path_containing_images>")
sys.exit(1)
images_to_package(sys.argv[1]).write_to_file(APKG_FILE)
logging.info(f"Deck created at: {APKG_FILE}")
+16
View File
@@ -0,0 +1,16 @@
# File and Directory Constants
IMAGE_KEY="image"
APKG_FILE="out.apkg"
CONVERTED_DIR = "converted"
TEXT_OCR_FILE = "final.txt"
IMAGE_EXTENSIONS = ['.png', '.jpg', '.jpeg']
DECK_JSON_FILE = "output_deck.json"
# API Constants
API_KEY_ENV = "OPENAI_API_KEY"
CHAT_MODEL = "gpt-3.5-turbo"
# Error Messages
NO_IMAGE_PART_ERROR = 'No image part'
NO_SELECTED_FILE_ERROR = 'No selected file'
INVALID_FILENAME_ERROR = 'Invalid filename'
-49
View File
@@ -1,49 +0,0 @@
import csv
import genanki
import sys
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607392319,
"Simple Model",
fields=[
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "Card 1",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def csv_to_anki(csv_path, output_path):
with open(csv_path, 'r', encoding='utf-8') as f:
reader = csv.reader(f)
# Skipping the header row
next(reader, None)
my_deck = genanki.Deck(2059400110, "CSV Deck")
for row in reader:
# Use row directly without splitting
question = row[0]
answer = ",".join(row[1:])
note = genanki.Note(
model=MY_MODEL,
fields=[question, answer]
)
my_deck.add_note(note)
genanki.Package(my_deck).write_to_file(output_path)
if __name__ == "__main__":
if len(sys.argv) != 3:
print("Usage: python convert.py <input_csv> <output_apkg>")
sys.exit(1)
input_csv = sys.argv[1]
output_apkg = sys.argv[2]
csv_to_anki(input_csv, output_apkg)
print(f"Deck created at: {output_apkg}")
+133
View File
@@ -0,0 +1,133 @@
import openai
import os
import json
import genanki
from logging_config import setup_logging
from constants import API_KEY_ENV, CHAT_MODEL
setup_logging()
API_KEY = os.environ.get(API_KEY_ENV)
if not API_KEY:
raise ValueError("Please set the OPENAI_API_KEY environment variable.")
openai.api_key = API_KEY
PROMPT_TEMPLATE = """
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
1. Every card should have a title, a question on the front, and an answer on the back.
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
6. Each index card should home in on answering a distinct question.
7. Limit each index card answer to no more than three sentences for brevity and clarity.
Structure your output as:
```
Deck Title: [Title of the Deck]
Cards:
- Title: [Card Title 1]
Front: [Question 1]
Back: [Answer 1]
- Title: [Card Title 2]
Front: [Question 2]
Back: [Answer 2]
... continue in this pattern
```
Content for reference:
{content}
"""
def prompt_for_card_content(text_content):
# Prepare the prompt
prompt = PROMPT_TEMPLATE.format(content=text_content)
# Get completion from the OpenAI ChatGPT API
response = openai.ChatCompletion.create(
model=CHAT_MODEL,
messages=[
{"role": "user", "content": prompt}
],
temperature=0,
)
# Extract content from response and save to a new file
return response.choices[0]['message']['content']
def response_to_json(response_text):
lines = [line.strip() for line in response_text.split("\n") if line.strip()]
deck_title = None
cards = []
current_card = {}
for line in lines:
if "Deck Title:" in line and not deck_title:
deck_title = line.split("Deck Title:", 1)[1].strip()
elif "Title:" in line:
if current_card: # If there's a card being processed, add it to cards
cards.append(current_card)
current_card = {}
current_card["Title"] = line.split("Title:", 1)[1].strip()
elif "Front:" in line:
current_card["Question"] = line.split("Front:", 1)[1].strip()
elif "Back:" in line:
current_card["Answer"] = line.split("Back:", 1)[1].strip()
if current_card: # Add the last card if it exists
cards.append(current_card)
return {
"DeckTitle": deck_title,
"Cards": cards
}
# Create a new model for our cards. This is necessary for genanki.
MY_MODEL = genanki.Model(
1607372319,
"Simple Model",
fields=[
{"name": "Title"},
{"name": "Question"},
{"name": "Answer"},
],
templates=[
{
"name": "{{Title}}",
"qfmt": "{{Question}}",
"afmt": "{{FrontSide}}<hr id='answer'>{{Answer}}",
},
])
def json_file_to_package(json_path):
with open(json_path, 'r', encoding='utf-8') as f:
json_data = json.load(f)
package = to_package(json_data)
return package
def to_package(deck_json):
deck_title = deck_json["DeckTitle"]
deck = genanki.Deck(1607372319, deck_title)
for card_json in deck_json["Cards"]:
title = card_json["Title"]
question = card_json["Question"]
answer = card_json["Answer"]
note = genanki.Note(
model=MY_MODEL,
fields=[title, question, answer]
)
deck.add_note(note)
return genanki.Package(deck)
+106
View File
@@ -0,0 +1,106 @@
import os
import sys
import logging
from logging_config import setup_logging
from subprocess import run, CalledProcessError
from concurrent.futures import ThreadPoolExecutor
from constants import CONVERTED_DIR, TEXT_OCR_FILE, IMAGE_EXTENSIONS
setup_logging()
def is_image_file(path):
return any(path.lower().endswith(ext) for ext in IMAGE_EXTENSIONS)
def ensure_directory_exists(directory):
if not os.path.exists(directory):
os.mkdir(directory)
def convert_image(image_path):
logging.info(f"Converting {image_path}...")
converted_path = os.path.join(CONVERTED_DIR, os.path.basename(image_path))
cmd = [
"convert",
image_path,
"-colorspace", "Gray",
"-resize", "300%",
"-threshold", "55%",
"-type", "Grayscale",
converted_path
]
try:
run(cmd, check=True)
logging.info(f"Converted image output to {converted_path}!")
return converted_path
except CalledProcessError:
logging.info(f"Error converting {image_path} with ImageMagick. Using original for Tesseract.")
return image_path
def ocr_image(image_path):
logging.info(f"OCR'ing {image_path}...")
base_name = os.path.basename(image_path)
root_name, _ = os.path.splitext(base_name)
text_filename = f"{root_name}.txt"
text_path = os.path.join(CONVERTED_DIR, text_filename)
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
try:
run(cmd, check=True)
logging.info(f"OCRed to {text_path}!")
return text_path
except CalledProcessError:
logging.info(f"Error processing {image_path} with Tesseract. Skipping.")
return None
def process_image(image_path):
converted_path = convert_image(image_path)
logging.info(f"OCR'ing image {image_path} (now at {converted_path})...")
text_path = ocr_image(converted_path)
if text_path and os.path.exists(text_path):
with open(text_path, 'r') as text_file:
text_content = text_file.read()
logging.info(f"Added text from {text_path} to final output.")
return text_content
else:
logging.info(f"Cannot locate {text_path}! Cannot add text to final output!")
return None
def process_images(directory_path):
final_text = []
ensure_directory_exists(CONVERTED_DIR)
image_paths = []
for root, dirs, files in os.walk(directory_path):
for file in files:
image_path = os.path.join(root, file)
if is_image_file(image_path):
image_paths.append(image_path)
# Use a ThreadPoolExecutor to process images in parallel
with ThreadPoolExecutor() as executor:
final_text = list(executor.map(process_image, image_paths))
# Filter out any None values and write the text to final.txt
final_text = [text for text in final_text if text is not None]
with open(TEXT_OCR_FILE, 'w') as f:
f.write("\n".join(final_text))
logging.info(f"All images processed! Final output saved to {TEXT_OCR_FILE}")
return final_text # Add this line
if __name__ == "__main__":
if len(sys.argv) != 2:
print("Usage: python images2text.py <directory_path>")
sys.exit(1)
process_images(sys.argv[1])
+11
View File
@@ -0,0 +1,11 @@
import logging
def setup_logging():
logging.basicConfig(level=logging.DEBUG,
format='%(asctime)s [%(levelname)s] - %(module)s: %(message)s',
datefmt='%Y-%m-%d %H:%M:%S')
# If you also want to save logs to a file, you can add the below lines.
# file_handler = logging.FileHandler('ankiai.log')
# file_handler.setFormatter(logging.Formatter('%(asctime)s [%(levelname)s] - %(module)s: %(message)s'))
# logging.getLogger().addHandler(file_handler)
+3
View File
@@ -1 +1,4 @@
genanki==0.8.0
Pillow==10.0.1
openai==0.28.0
Flask==2.3.3
+54
View File
@@ -0,0 +1,54 @@
import os
import tempfile
import shutil
import logging
from flask import Flask, request, send_from_directory, jsonify
from werkzeug.utils import secure_filename
from ankiai import images_to_package
from constants import IMAGE_KEY, APKG_FILE, NO_IMAGE_PART_ERROR, NO_SELECTED_FILE_ERROR, INVALID_FILENAME_ERROR
from logging_config import setup_logging
setup_logging()
app = Flask(__name__)
def save_uploaded_images(images, directory):
for img in images:
safe_filename = secure_filename(img.filename)
if not safe_filename:
raise ValueError(INVALID_FILENAME_ERROR)
filename = os.path.join(directory, safe_filename)
img.save(filename)
@app.route('/deck-from-images', methods=['POST'])
def deck_from_images():
if IMAGE_KEY not in request.files:
return jsonify({'error': NO_IMAGE_PART_ERROR}), 400
images = request.files.getlist(IMAGE_KEY)
if not images or not any(img.filename != '' for img in images):
return jsonify({'error': NO_SELECTED_FILE_ERROR}), 400
temp_dir = tempfile.mkdtemp()
save_uploaded_images(images, temp_dir)
try:
images_to_package(temp_dir).write_to_file(APKG_FILE)
logging.info(f"Anki package written to {APKG_FILE}")
return send_from_directory('.', APKG_FILE, as_attachment=True)
except Exception as e:
logging.error("Exception occurred: "+str(e), exc_info=True)
return jsonify({'error': str(e)}), 500
finally:
shutil.rmtree(temp_dir)
if __name__ == '__main__':
app.run(debug=True)