Compare commits
6
Commits
e96f23ddc8
...
master
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9d33c2ee10 | ||
|
|
aa17dcc48a | ||
|
|
6c1ece44ab | ||
|
|
3f3cc276f2 | ||
|
|
34511cca74 | ||
|
|
963d99404e |
@@ -20,16 +20,27 @@ AnkiAI is a tool that leverages OCR (Optical Character Recognition) and GPT-3's
|
||||
|
||||
### Requirements
|
||||
|
||||
To run AnkiAI, you'll need to have the following dependencies installed:
|
||||
#### ImageMagick
|
||||
|
||||
```
|
||||
genanki==0.8.0
|
||||
Pillow
|
||||
openai
|
||||
flask
|
||||
ImageMagick is a software suite that allows you to create, edit, and compose bitmap images. It can read, convert, and write images in a variety of formats (over 100) including DPX, EXR, GIF, JPEG, JPEG-2000, PDF, PhotoCD, PNG, Postscript, SVG, and TIFF. In the AnkiAI project, it is used for preprocessing images to improve the performance of OCR.
|
||||
|
||||
```bash
|
||||
sudo apt-get update
|
||||
sudo apt-get install imagemagick
|
||||
```
|
||||
|
||||
You can install these via `pip` using the `requirements.txt` file:
|
||||
#### Tesseract
|
||||
|
||||
You need Tesseract for the OCR functionality:
|
||||
|
||||
```bash
|
||||
sudo apt-get install tesseract-ocr
|
||||
```
|
||||
### Python Dependencies
|
||||
|
||||
To ensure consistent functionality, it's crucial to use the provided `requirements.txt` file which pins dependencies to known compatible versions.
|
||||
|
||||
You can install the Python dependencies via `pip` using the `requirements.txt` file:
|
||||
|
||||
```bash
|
||||
pip install -r requirements.txt
|
||||
@@ -38,7 +49,11 @@ pip install -r requirements.txt
|
||||
### How to Run
|
||||
|
||||
1. **Environment Variables**: Make sure to set the `OPENAI_API_KEY` environment variable to your OpenAI API key.
|
||||
|
||||
|
||||
```bash
|
||||
export OPENAI_API_KEY=sk-myapikey
|
||||
```
|
||||
|
||||
2. **Run the Flask server**:
|
||||
|
||||
```bash
|
||||
@@ -55,6 +70,29 @@ pip install -r requirements.txt
|
||||
python ankiai.py <directory_path_containing_images>
|
||||
```
|
||||
|
||||
### Example curl commands to interact with the service:
|
||||
|
||||
You can make POST requests to the server using curl. Here are some examples from the command line history:
|
||||
|
||||
```bash
|
||||
curl -X POST -o deck.apkg \
|
||||
-F "image=@/home/ubuntu/Pictures/image1.png" \
|
||||
-F "image=@/home/ubuntu/Pictures/image2.png" \
|
||||
-F "image=@/home/ubuntu/Pictures/image3.png" \
|
||||
http://localhost:5000/deck-from-images
|
||||
```
|
||||
|
||||
Batch processing of images:
|
||||
|
||||
```bash
|
||||
for file in /home/ubuntu/Pictures/*; do
|
||||
if [[ -f "$file" ]]; then
|
||||
basefile=$(basename "$file");
|
||||
curl -X POST -o "deck-${basefile}.apkg" -F "image=@${file}" http://localhost:5000/deck-from-images;
|
||||
fi;
|
||||
done
|
||||
```
|
||||
|
||||
### How to Debug (VSCode Users)
|
||||
|
||||
- Open the project in VSCode.
|
||||
|
||||
+19
-16
@@ -15,28 +15,31 @@ if not API_KEY:
|
||||
|
||||
openai.api_key = API_KEY
|
||||
|
||||
# Given prompt template
|
||||
PROMPT_TEMPLATE = """
|
||||
Please come up with a title for the deck and a set of 10 index cards for memorization,
|
||||
including a title, front, and back for each card. The index cards should completely
|
||||
capture the main points and themes of the text. In addition, they should contain any
|
||||
numbers or data that humans might find difficult to remember. The goal of the index
|
||||
card set is that one who memorizes it can provide a summary of the text to someone
|
||||
else, conveying the main points and themes.
|
||||
Please craft a title for the deck and generate a comprehensive set of index cards based on the provided text. Follow these guidelines:
|
||||
|
||||
You will provide the deck title, and the titles, questions, and answers for each card
|
||||
in a structured format as follows:
|
||||
1. Every card should have a title, a question on the front, and an answer on the back.
|
||||
2. Each answer must contain at least one concrete fact that is not evident from its corresponding question.
|
||||
3. Ensure inclusion of numbers, data, or intricate details that would be challenging for individuals to remember.
|
||||
4. The goal is to enable someone who learns this set to competently convey both the overarching themes and intricate details of the text to another person.
|
||||
5. Create one index card for every 2-4 sentences of the content. The exact number depends on the density of the information. Aim for completeness over brevity.
|
||||
6. Each index card should home in on answering a distinct question.
|
||||
7. Limit each index card answer to no more than three sentences for brevity and clarity.
|
||||
|
||||
Structure your output as:
|
||||
```
|
||||
Deck Title: Title of the Deck
|
||||
Deck Title: [Title of the Deck]
|
||||
Cards:
|
||||
- Title: Card Title 1
|
||||
Front: What is the capital of New York?
|
||||
Back: Albany
|
||||
- Title: Card Title 2
|
||||
Front: Where in the world is Carmen San Diego?
|
||||
Back: Nobody knows
|
||||
- Title: [Card Title 1]
|
||||
Front: [Question 1]
|
||||
Back: [Answer 1]
|
||||
- Title: [Card Title 2]
|
||||
Front: [Question 2]
|
||||
Back: [Answer 2]
|
||||
... continue in this pattern
|
||||
```
|
||||
|
||||
Content for reference:
|
||||
{content}
|
||||
"""
|
||||
|
||||
|
||||
+5
-1
@@ -44,7 +44,11 @@ def convert_image(image_path):
|
||||
|
||||
def ocr_image(image_path):
|
||||
logging.info(f"OCR'ing {image_path}...")
|
||||
text_filename = os.path.basename(image_path).replace(".jpg", ".txt")
|
||||
|
||||
base_name = os.path.basename(image_path)
|
||||
root_name, _ = os.path.splitext(base_name)
|
||||
text_filename = f"{root_name}.txt"
|
||||
|
||||
text_path = os.path.join(CONVERTED_DIR, text_filename)
|
||||
cmd = ["tesseract", image_path, text_path.replace(".txt", "")]
|
||||
try:
|
||||
|
||||
+3
-3
@@ -1,4 +1,4 @@
|
||||
genanki==0.8.0
|
||||
Pillow
|
||||
openai
|
||||
flask
|
||||
Pillow==10.0.1
|
||||
openai==0.28.0
|
||||
Flask==2.3.3
|
||||
Reference in New Issue
Block a user