Usando o Amazon Textract como OCR para Extração de Dados no DynamoDB

Last update: Jan 19, 2022

Overview

dio-live-textract2

Repositório de código para o live coding do dia 05/10/2021 sobre extração de dados estruturados e gravação em banco de dados a partir do Amazon Textract.

Serviços utilizados

Amazon Textract
AWS Lambda
Amazon S3
Amazon DynamoDB

Desenvolvimento

Criando um bucket no Amazon S3

S3 Console -> Create bucket -> Bucket name "dio-live-input-data" -> Manter as configurações padrão -> Create bucket

Processando imagens no Amazon Textract

Textract Console -> Select Document -> Analyze Document -> Tables
Download results -> Salvar arquivo .zip

Criando uma tabela no DynamoDB

DynamoDB Console -> Tables -> Create Table -> Partition key "cod" -> Create table

Implementando a função lambda

Lambda Console -> Functions -> Create function
Use a blueprint -> "s3-get-object-python"
Function name "dio-live-csv-to-db"
Execution role -> "Create a new role from AWS policy templates" -> Role name "S3ToDynamoDBRole"
S3 Trigger -> Bucket criado anteriormente
Create function
Substituir o código gerado pelo código da pasta /src deste repositório (Obs: atenção para o nome da tabela, deve ser substituído pelo nome da sua)

Passo adicional: Criando um layer com a biblioteca boto3 do Python

Lambda Console -> Additional Resources -> Layers
Name "boto3_layer" -> Upload a .zip file -> baixe e insira o arquivo .zip contido na pasta /src deste respositório
Compatible architecture "x86_64"
Compatible runtimes "Python3.7" (É necessário ser Python3.7 para ser compatível com a versão do blueprint utilizado)
Create
Na função lambda criada -> Selecione layers no diagrama -> Add layer -> Custom layers "boto3_layer" -> Version 1 -> Add

Configurando permissões no Lambda para o DynamoDB

Lambda Console -> Functions -> Selecione a função criada -> Configuration -> Permission -> Execution Role -> Abrir a role criada no Amazon IAM
No IAM -> Permission -> Add inline policy -> Choose a service "DynamoDB" -> Write "PutItem"
Resources -> Selecionar o Arn da sua tabela -> Selecionar a sua região -> Add -> Review Policy -> Name "LambdaDynamoDBPolicy" -> Create policy

Utilizando a aplicação

No Amazon Textract

Amazon Textract Console -> Select Document -> Choose file -> Buscar o arquivo a ser analisado
Download results

No Amazon S3

Extrair o arquivo table_1.csv do arquivo baixado do Amazon Textract
Acessar o bucket criado anteriormente -> Upload -> Selecionar o arquivo table_1.csv -> Upload

No DynamoDB

Tables -> Acessar a tabela criada -> View Items

Usando o Amazon Textract como OCR para Extração de Dados no DynamoDB

Related tags

Overview

dio-live-textract2

Serviços utilizados

Desenvolvimento

Criando um bucket no Amazon S3

Processando imagens no Amazon Textract

Criando uma tabela no DynamoDB

Implementando a função lambda

Passo adicional: Criando um layer com a biblioteca boto3 do Python

Configurando permissões no Lambda para o DynamoDB

Utilizando a aplicação

No Amazon Textract

No Amazon S3

No DynamoDB

Owner

hugoportela

A python scripts that uses 3 different feature extraction methods such as SIFT, SURF and ORB to find a book in a video clip and project trailer of a movie based on that book, on to it.

Indonesian ID Card OCR using tesseract OCR

A simple document layout analysis using Python-OpenCV

Text modding tools for FF7R (Final Fantasy VII Remake)

An OCR evaluation tool

Fun program to overlay a mask to yourself using a webcam

Satoshi is a discord bot template in python using discord.py that allow you to track some live crypto prices with your own discord bot.

Here use convulation with sobel filter from scratch in opencv python .

Source code of RRPN ---- Arbitrary-Oriented Scene Text Detection via Rotation Proposals

A synthetic data generator for text recognition

MXNet OCR implementation. Including text recognition and detection.

Image Detector and Convertor App created using python's Pillow, OpenCV, cvlib, numpy and streamlit packages.

A program that takes in the hand gesture displayed by the user and translates ASL.

Convert Text-to Handwriting Using Python

This repository provides train＆test code, dataset, det.&rec. annotation, evaluation script, annotation tool, and ranking.

A python screen recorder for low-end computers, provides high quality video output.

📷 Face Recognition using Haar-Cascade Classifier, OpenCV, and Python

Qrcode Attendence System with Opencv and Pyzbar

[EMNLP 2021] Improving and Simplifying Pattern Exploiting Training

Official code for "Bridging Video-text Retrieval with Multiple Choice Questions", CVPR 2022 (Oral).