Go to file

Julius Unverfehrt e7b28f5bda Pull request #18 : Remove pil

Merge in RR/cv-analysis from remove_pil to master

Squashed commit of the following:

commit 83c8d88f3d48404251470176c70979ee75ae068b
Author: Julius Unverfehrt <julius.unverfehrt@iqser.com>
Date:   Thu Jul 21 10:51:51 2022 +0200

    remove deprecated server tests

commit cebc03b5399ac257a74036b41997201f882f5b74
Author: Julius Unverfehrt <julius.unverfehrt@iqser.com>
Date:   Thu Jul 21 10:51:08 2022 +0200

    remove deprecated server tests

commit ce2845b0c51f001b7b5b8b195d6bf7e034ec4e39
Author: Julius Unverfehrt <julius.unverfehrt@iqser.com>
Date:   Wed Jul 20 17:05:00 2022 +0200

    repair tests to work without pillow WIP

commit 023fdab8322f28359a24c63e32635a3d0deccbe4
Author: Isaac Riley <Isaac.Riley@iqser.com>
Date:   Wed Jul 20 16:40:36 2022 +0200

    fixed typo

commit 33850ca83a175f74789ae6b9bebd057ed84b7fb3
Author: Isaac Riley <Isaac.Riley@iqser.com>
Date:   Wed Jul 20 16:38:37 2022 +0200

    fixed import from refactored open_img.py

commit dbc6d345f074e538948e2c4f94ebed8a5ef520bc
Author: Isaac Riley <Isaac.Riley@iqser.com>
Date:   Wed Jul 20 16:32:42 2022 +0200

    removed PIL from production code, now inly in scripts

2022-07-21 13:25:00 +02:00

.dvc

Pull request #4 : Restructuring and renaming of module

2022-02-05 16:14:24 +01:00

bamboo-specs

update dependencies

2022-06-23 16:54:13 +02:00

cv_analysis

Pull request #18 : Remove pil

2022-07-21 13:25:00 +02:00

data

Pull request #16 : Add table parsing fixtures

2022-07-11 12:25:16 +02:00

incl

Pull request #18 : Remove pil

2022-07-21 13:25:00 +02:00

scripts

Pull request #18 : Remove pil

2022-07-21 13:25:00 +02:00

setup

change name from vidocp to cv-analysis

2022-03-23 13:46:57 +01:00

src

Pull request #15 : Refactor logging

2022-07-11 09:36:57 +02:00

test

Pull request #18 : Remove pil

2022-07-21 13:25:00 +02:00

.coveragerc

Pull request #11 : Integrate new pyinfra

2022-06-23 14:45:08 +02:00

.dockerignore

add new files for containerization; still some work to do, but want to merge in tests first

2022-03-02 09:38:56 +01:00

.dvcignore

Pull request #4 : Restructuring and renaming of module

2022-02-05 16:14:24 +01:00

.gitignore

Pull request #16 : Add table parsing fixtures

2022-07-11 12:25:16 +02:00

.gitmodules

Pull request #11 : Integrate new pyinfra

2022-06-23 14:45:08 +02:00

config.yaml

Pull request #15 : Refactor logging

2022-07-11 09:36:57 +02:00

Dockerfile

Pull request #13 : Add pdf coord conversion

2022-07-07 11:35:12 +02:00

Dockerfile_base

Pull request #11 : Integrate new pyinfra

2022-06-23 14:45:08 +02:00

pytest.ini

Pull request #11 : Integrate new pyinfra

2022-06-23 14:45:08 +02:00

README.md

tiny change to test build server

2022-03-23 14:35:00 +01:00

requirements.txt

Pull request #17 : Add pdf2array func

2022-07-20 11:01:55 +02:00

setup.py

change name from vidocp to cv-analysis

2022-03-23 13:46:57 +01:00

sonar-project.properties

fixed build config minutia

2022-03-22 14:06:15 +01:00

README.md

cv-analysis — Visual (CV-Based) Document Parsing

This repository implements computer vision based approaches for detecting and parsing visual features such as tables or previous redactions in documents.

Installation

git clone ssh://git@git.iqser.com:2222/rr/cv-analysis.git
cd cv-analysis

python -m venv env
source env/bin/activate

pip install -e .
pip install -r requirements.txt

dvc pull

Usage

As an API

The module provided functions for the individual tasks that all return some kind of collection of points, depending on the specific task.

Redaction Detection (API)

The below snippet shows hot to find the outlines of previous redactions.

from cv_analysis.redaction_detection import find_redactions
import pdf2image 
import numpy as np


pdf_path = ...
page_index = ...

page = pdf2image.convert_from_path(pdf_path, first_page=page_index, last_page=page_index)[0]
page = np.array(page)

redaction_contours = find_redactions(page)

As a CLI Tool

Core API functionalities can be used through a CLI.

Table Parsing

The tables parsing utility detects and segments tables into individual cells.

python scripts/annotate.py data/test_pdf.pdf 7 --type table

The below image shows a parsed table, where each table cell has been detected individually.

Redaction Detection (CLI)

The redaction detection utility detects previous redactions in PDFs (filled black rectangles).

python scripts/annotate.py data/test_pdf.pdf 2 --type redaction

The below image shows the detected redactions with green outlines.

Layout Parsing

The layout parsing utility detects elements such as paragraphs, tables and figures.

python scripts/annotate.py data/test_pdf.pdf 7 --type layout

The below image shows the detected layout elements on a page.

Figure Detection

The figure detection utility detects figures specifically, which can be missed by the generic layout parsing utility.

python scripts/annotate.py data/test_pdf.pdf 3 --type figure

The below image shows the detected figure on a page.

Running as a service

Building

Build base image

bash setup/docker.sh

Build head image

docker build -f Dockerfile -t cv-analysis . --build-arg BASE_ROOT=""

Usage (service)

Shell 1

docker run --rm --net=host --rm cv-analysis

Shell 2

python scripts/client_mock.py --pdf_path /path/to/a/pdf

Releases 37

Release 2.29.0 Latest

2025-01-16 09:31:10 +01:00

Languages

Python 91.1%

Shell 3%

Makefile 2.4%

Dockerfile 2.3%

Nix 1.2%