Compare commits
6
Commits
e96f23ddc8
..
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9d33c2ee10 | ||
|
|
aa17dcc48a | ||
|
|
6c1ece44ab | ||
|
|
3f3cc276f2 | ||
|
|
34511cca74 | ||
|
|
963d99404e |
@@ -20,16 +20,27 @@ AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's
|
|||||||
|
|
||||||
### Requirements
|
### Requirements
|
||||||
|
|
||||||
To run AnkiAI, you'll need to have the following dependencies installed:
|
#### ImageMagick
|
||||||
|
|
||||||
```
|
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
|
||||||
genanki==0.8.0
|
|
||||||
Pillow
|
```bash
|
||||||
openai
|
sudo apt-get update
|
||||||
flask
|
sudo apt-get install imagemagick
|
||||||
```
|
```
|
||||||
|
|
||||||
You can install these via `pip` using the `requirements.txt` file:
|
#### Tesseract
|
||||||
|
|
||||||
|
You need Tesseract for the OCR functionality:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
sudo apt-get install tesseract-ocr
|
||||||
|
```
|
||||||
|
### Python Dependencies
|
||||||
|
|
||||||
|
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
|
||||||
|
|
||||||
|
You can install the Python dependencies via `pip` using the `requirements.txt` file:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install -r requirements.txt
|
pip install -r requirements.txt
|
||||||
@@ -38,7 +49,11 @@ pip install -r requirements.txt
|
|||||||
### How to Run
|
### How to Run
|
||||||
|
|
||||||
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
|
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
export OPENAI_API_KEY=sk-myapikey
|
||||||
|
```
|
||||||
|
|
||||||
2. **Run the Flask server**:
|
2. **Run the Flask server**:
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
@@ -55,6 +70,29 @@ pip install -r requirements.txt
|
|||||||
python ankiai.py <directory_path_containing_images>
|
python ankiai.py <directory_path_containing_images>
|
||||||
```
|
```
|
||||||
|
|
||||||
|
### Example curl commands to interact with the service:
|
||||||
|
|
||||||
|
You can make POST requests to the server using curl. Here are some examples from the command line history:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -X POST -o deck.apkg \
|
||||||
|
-F "image=@/home/ubuntu/Pictures/image1.png" \
|
||||||
|
-F "image=@/home/ubuntu/Pictures/image2.png" \
|
||||||
|
-F "image=@/home/ubuntu/Pictures/image3.png" \
|
||||||
|
http://localhost:5000/deck-from-images
|
||||||
|
```
|
||||||
|
|
||||||
|
Batch processing of images:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
for file in /home/ubuntu/Pictures/*; do
|
||||||
|
if [[ -f "$file" ]]; then
|
||||||
|
basefile=$(basename "$file");
|
||||||
|
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
|
||||||
|
fi;
|
||||||
|
done
|
||||||
|
```
|
||||||
|
|
||||||
### How to Debug (VSCode Users)
|
### How to Debug (VSCode Users)
|
||||||
|
|
||||||
- Open the project in VSCode.
|
- Open the project in VSCode.
|
||||||
|
|||||||
+19
-16
@@ -15,28 +15,31 @@ if not API_KEY:
|
|||||||
|
|
||||||
openai.api_key = API_KEY
|
openai.api_key = API_KEY
|
||||||
|
|
||||||
# Given prompt template
|
|
||||||
PROMPT_TEMPLATE = """
|
PROMPT_TEMPLATE = """
|
||||||
Please come up with a title for the deck and a set of 10 index cards for memorization,
|
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
|
||||||
including a title, front, and back for each card. The index cards should completely
|
|
||||||
capture the main points and themes of the text. In addition, they should contain any
|
|
||||||
numbers or data that humans might find difficult to remember. The goal of the index
|
|
||||||
card set is that one who memorizes it can provide a summary of the text to someone
|
|
||||||
else, conveying the main points and themes.
|
|
||||||
|
|
||||||
You will provide the deck title, and the titles, questions, and answers for each card
|
1. Every card should have a title, a question on the front, and an answer on the back.
|
||||||
in a structured format as follows:
|
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
|
||||||
|
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
|
||||||
|
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
|
||||||
|
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
|
||||||
|
6. Each index card should home in on answering a distinct question.
|
||||||
|
7. Limit each index card answer to no more than three sentences for brevity and clarity.
|
||||||
|
|
||||||
|
Structure your output as:
|
||||||
```
|
```
|
||||||
Deck Title: Title of the Deck
|
Deck Title: [Title of the Deck]
|
||||||
Cards:
|
Cards:
|
||||||
- Title: Card Title 1
|
- Title: [Card Title 1]
|
||||||
Front: What is the capital of New York?
|
Front: [Question 1]
|
||||||
Back: Albany
|
Back: [Answer 1]
|
||||||
- Title: Card Title 2
|
- Title: [Card Title 2]
|
||||||
Front: Where in the world is Carmen San Diego?
|
Front: [Question 2]
|
||||||
Back: Nobody knows
|
Back: [Answer 2]
|
||||||
|
... continue in this pattern
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Content for reference:
|
||||||
{content}
|
{content}
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
|||||||
+5
-1
@@ -44,7 +44,11 @@ def convert_image(image_path):
|
|||||||
|
|
||||||
def ocr_image(image_path):
|
def ocr_image(image_path):
|
||||||
logging.info(f"OCR'ing {image_path}...")
|
logging.info(f"OCR'ing {image_path}...")
|
||||||
text_filename = os.path.basename(image_path).replace(".jpg", ".txt")
|
|
||||||
|
base_name = os.path.basename(image_path)
|
||||||
|
root_name, _ = os.path.splitext(base_name)
|
||||||
|
text_filename = f"{root_name}.txt"
|
||||||
|
|
||||||
text_path = os.path.join(CONVERTED_DIR, text_filename)
|
text_path = os.path.join(CONVERTED_DIR, text_filename)
|
||||||
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
|
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
|
||||||
try:
|
try:
|
||||||
|
|||||||
+3
-3
@@ -1,4 +1,4 @@
|
|||||||
genanki==0.8.0
|
genanki==0.8.0
|
||||||
Pillow
|
Pillow==10.0.1
|
||||||
openai
|
openai==0.28.0
|
||||||
flask
|
Flask==2.3.3
|
||||||
Reference in New Issue
Block a user