Compare commits

...
40 Commits
Author SHA1 Message Date
deiflaender 0c8ff214e4 test serializer error 2022-06-14 14:25:27 +02:00
deiflaender 2a6ee02f9f RED-4129: Fixed calculation of imported redaction intersections for all rotation cases 2022-06-14 12:10:25 +02:00
Dominique Eiflaender 7dad091e19 Pull request #409: RED-3805: Added transpacency to images
Merge in RED/redaction-service from RED-3805 to master

* commit 'f7b584617bd5bc7b2d71e44c407b0d519c1f5a75':
  RED-3805: Added transpacency to images
2022-06-10 13:27:16 +02:00
deiflaender f7b584617b RED-3805: Added transpacency to images 2022-06-10 13:24:14 +02:00
Dominique Eiflaender 02a264d586 Pull request #408: RED-4114: Added all missing cases for text rotations
Merge in RED/redaction-service from RED-4114 to master

* commit 'dacd7b725d1bcdd17193aaefa045bd9555b8528a':
  RED-4114: Added all missing cases for text rotations
2022-06-10 12:20:16 +02:00
deiflaender dacd7b725d RED-4114: Added all missing cases for text rotations 2022-06-10 12:15:53 +02:00
Dominique Eiflaender 7ee338a067 Pull request #407: RED-4176: Only apply APPROVED resizes in to annotations in analysis
Merge in RED/redaction-service from RED-4176 to master

* commit '19d48e207b6fb4169adb894ced089fb3c2f9c49f':
  RED-4176: Only apply APPROVED resizes in to annotations in analysis
2022-06-09 12:11:28 +02:00
deiflaender 19d48e207b RED-4176: Only apply APPROVED resizes in to annotations in analysis 2022-06-09 12:08:21 +02:00
Dominique Eiflaender e22b625e9c Pull request #406: RED-4114: Fixed text rotation for rotation = 0 and text direction = 180
Merge in RED/redaction-service from RED-4114 to master

* commit '95a2d71af0a69c06a40e4d39d1f9f1286d726028':
  RED-4114: Fixed text rotation for rotation = 0 and text direction = 180
2022-06-09 09:58:46 +02:00
deiflaender 95a2d71af0 RED-4114: Fixed text rotation for rotation = 0 and text direction = 180 2022-06-09 09:55:46 +02:00
Corina Olariu 6fec4cc209 Pull request #404: RED-4029 - Do not log on ERROR level if something happens that is expected
Merge in RED/redaction-service from RED_4029_1 to master

* commit 'fa3cbc938c6d5066b4c6706bc0c98ac030f4c498':
  RED-4029 - Do not log on ERROR level if something happens that is expected - put back null pointer exception to error logging
2022-06-07 10:49:18 +02:00
devplant fa3cbc938c RED-4029 - Do not log on ERROR level if something happens that is expected
- put back null pointer exception to error logging
2022-06-07 11:39:55 +03:00
Dominique Eiflaender f6b4d2a067 Pull request #403: RED-4014: Renamed wrong metrics name
Merge in RED/redaction-service from RED-4014 to master

* commit '0967e8e64a99654733df1e860d297e03587c7dea':
  RED-4014: Renamed wrong metrics name
2022-06-07 10:27:23 +02:00
deiflaender 0967e8e64a RED-4014: Renamed wrong metrics name 2022-06-07 10:14:59 +02:00
Dominique Eiflaender 20c7e19343 Pull request #402: RED-4014: Renamed and added metrics
Merge in RED/redaction-service from RED-4014 to master

* commit '6fb7600de2ab7a00d6071de3c701c3a7410e8fb7':
  RED-4014: Renamed and added metrics
2022-06-07 09:33:06 +02:00
deiflaender 6fb7600de2 RED-4014: Renamed and added metrics 2022-06-07 09:28:38 +02:00
Corina Olariu fbd55c15ef Pull request #401: RED-4029 - Do not log on ERROR level if something happens that is expected
Merge in RED/redaction-service from RED-4029 to master

* commit 'ffa251643a2e4d711bac01469d31e09c3c8b8290':
  RED-4029 - Do not log on ERROR level if something happens that is expected - replace error logs
2022-06-06 20:50:27 +02:00
devplant ffa251643a RED-4029 - Do not log on ERROR level if something happens that is expected
- replace error logs
2022-06-03 15:31:00 +03:00
Dominique Eiflaender ac19defa1c Pull request #400: RED-4064: Add resize redactions sections to reanalysis sections
Merge in RED/redaction-service from RED-4064 to master

* commit '21ca25f3309f0706c1d708eb77724824ef10168d':
  RED-4064: Add resize redactions sections to reanalysis sections
2022-06-03 12:09:17 +02:00
deiflaender 21ca25f330 RED-4064: Add resize redactions sections to reanalysis sections 2022-06-03 12:01:05 +02:00
Dominique Eiflaender e9b9fbb91b Pull request #399: RED-4064: Apply all resize redactions directly after find, calculate surrounding text in renalysis
Merge in RED/redaction-service from RED-4064 to master

* commit '3f76401c7d051ee100eab8a36847c2e8a5426900':
  RED-4064: Apply all resize redactions directly after find, calculate surrounding text in renalysis
2022-06-03 10:54:49 +02:00
deiflaender 3f76401c7d RED-4064: Apply all resize redactions directly after find, calculate surrounding text in renalysis 2022-06-03 10:51:21 +02:00
Dominique Eiflaender a16f9c6d6e Pull request #398: RED-4064: Fixed override logic for resized entries
Merge in RED/redaction-service from RED-4064 to master

* commit '7ca3416d1d809189430540603034b67dfddd5bcd':
  RED-4064: Fixed override logic for resized entries
2022-06-02 12:43:07 +02:00
deiflaender 7ca3416d1d RED-4064: Fixed override logic for resized entries 2022-06-02 12:40:17 +02:00
Dominique Eiflaender 85db826f89 Pull request #397: RED-4064: Apply resize redactions also in redactCell
Merge in RED/redaction-service from RED-4064 to master

* commit '6fac40fd8dea79c5575cee2b480269d43baaef7a':
  RED-4064: Apply resize redactions also in redactCell
2022-06-02 11:30:46 +02:00
deiflaender 6fac40fd8d RED-4064: Apply resize redactions also in redactCell 2022-06-02 11:27:26 +02:00
Dominique Eiflaender 23fe93630b Pull request #396: RED-4064: Fixed find dictionary entries at position where rules entries are resized to smaller entries
Merge in RED/redaction-service from RED-4064 to master

* commit '778c9756495f14ead45fadd8b906f4301e8b3ec7':
  RED-4064: Fixed find dictionary entries at position where rules entries are resized to smaller entries
2022-06-02 10:41:27 +02:00
deiflaender 778c975649 RED-4064: Fixed find dictionary entries at position where rules entries are resized to smaller entries 2022-06-02 10:35:38 +02:00
Philipp Schramm 50b2908763 Pull request #395: RED-3816: Bugfix with adding images to sections
Merge in RED/redaction-service from bugfix/RED-3816 to master

* commit 'c1192ceefd0717958badc27d2c2a65a8059952a4':
  RED-3816: Bugfix with adding images to sections
2022-05-30 14:43:35 +02:00
Philipp Schramm c1192ceefd RED-3816: Bugfix with adding images to sections 2022-05-30 13:52:57 +02:00
Timo Bejan 77584c9a5a Pull request #394: RED-3800 String Performance matching test
Merge in RED/redaction-service from RED-3800-string-performance-test to master

* commit '21d717f0837c9cc9b2f5b251d372db88b25a7b6d':
  RED-3800 String Performance matching test
  RED-3800 String Performance matching test
2022-05-24 11:09:08 +02:00
Timo Bejan 21d717f083 RED-3800 String Performance matching test 2022-05-24 12:01:10 +03:00
Timo Bejan c85ce25ed4 RED-3800 String Performance matching test 2022-05-24 11:58:34 +03:00
Philipp Schramm f608d991a3 Pull request #393: RED-3816: Implemented ThenAction to redact section with rectangle
Merge in RED/redaction-service from RED-3816 to master

* commit 'eb7137315208dadf3a679e21b05ec16d53c4ec43':
  RED-3816: Implemented ThenAction to redact section with rectangle
2022-05-24 09:00:07 +02:00
Ali Oezyetimoglu e914b09d1e Pull request #392: RED-3674: Fixed force redactions overrides idRemoval with different id
Merge in RED/redaction-service from RED-3674-rs1 to master

* commit '87d64b6d1285eaf7cd5ee3c2fb899f92063bc70f':
  RED-3674: Fixed force redactions overrides idRemoval with different id
2022-05-23 18:48:10 +02:00
aoezyetimoglu 87d64b6d12 RED-3674: Fixed force redactions overrides idRemoval with different id 2022-05-23 17:38:23 +02:00
Philipp Schramm eb71373152 Merge branch 'master' into RED-3816 2022-05-23 16:27:12 +02:00
Philipp Schramm a65e23e45d RED-3816: Implemented ThenAction to redact section with rectangle 2022-05-23 16:26:47 +02:00
Dominique Eiflaender 4aaef260d7 Pull request #390: RED-4075: Add legalBasis if Rule Entity overrides Ai Recommendation
Merge in RED/redaction-service from RED-4075 to master

* commit 'c669f9a09a3e5e254d41715f65fef3b158873338':
  RED-4075: Add legalBasis if Rule Entity overrides Ai Recommendation
2022-05-23 14:05:38 +02:00
deiflaender c669f9a09a RED-4075: Add legalBasis if Rule Entity overrides Ai Recommendation 2022-05-23 13:54:27 +02:00
44 changed files with 8539 additions and 368 deletions
@@ -187,7 +187,7 @@ public class PDFLinesTextStripper extends PDFTextStripper {
rulings.addAll(path);
}
} catch (UnsupportedOperationException e) {
log.error("UnsupportedOperationException: " + getGraphicsState().getStrokingColor()
log.debug("UnsupportedOperationException: " + getGraphicsState().getStrokingColor()
.getColorSpace()
.getName() + " or " + getGraphicsState().getNonStrokingColor()
.getColorSpace()
@@ -1871,7 +1871,7 @@ public class PDFTextStripper extends LegacyPDFStreamEngine
}
catch (IOException e)
{
LOG.error("Could not close BidiMirroring.txt ", e);
LOG.debug("Could not close BidiMirroring.txt ", e);
}
}
}
@@ -1,16 +1,22 @@
package com.iqser.red.service.redaction.v1.server.parsing.model;
import org.apache.pdfbox.text.TextPosition;
import org.springframework.beans.BeanUtils;
import com.dslplatform.json.CompiledJson;
import com.dslplatform.json.JsonAttribute;
import com.fasterxml.jackson.annotation.JsonIgnore;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
import lombok.SneakyThrows;
import org.apache.pdfbox.text.TextPosition;
import org.springframework.beans.BeanUtils;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
@CompiledJson
public class RedTextPosition {
@@ -1,24 +1,30 @@
package com.iqser.red.service.redaction.v1.server.parsing.model;
import java.util.ArrayList;
import java.util.List;
import java.util.stream.Collectors;
import org.apache.pdfbox.text.TextPosition;
import com.dslplatform.json.CompiledJson;
import com.dslplatform.json.JsonAttribute;
import com.fasterxml.jackson.annotation.JsonIgnore;
import com.fasterxml.jackson.annotation.JsonIgnoreProperties;
import com.iqser.red.service.redaction.v1.model.Point;
import com.iqser.red.service.redaction.v1.model.Rectangle;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.apache.pdfbox.text.TextPosition;
import java.util.ArrayList;
import java.util.List;
import java.util.stream.Collectors;
@Slf4j
@Data
@Builder
@CompiledJson
@NoArgsConstructor
@AllArgsConstructor
@JsonIgnoreProperties({"empty"})
public class TextPositionSequence implements CharSequence {
@@ -28,6 +34,7 @@ public class TextPositionSequence implements CharSequence {
private float x1;
private float x2;
public TextPositionSequence(int page) {
this.page = page;
@@ -255,7 +262,8 @@ public class TextPositionSequence implements CharSequence {
@JsonAttribute(ignore = true)
public Rectangle getRectangle() {
log.debug("Page: '{}', Word: '{}', Rotation: '{}', textRotation {}", page, toString(), textPositions.get(0).getRotation(), textPositions.get(0).getDir());
log.debug("Page: '{}', Word: '{}', Rotation: '{}', textRotation {}", page, toString(), textPositions.get(0)
.getRotation(), textPositions.get(0).getDir());
float height = getTextHeight();
@@ -264,26 +272,24 @@ public class TextPositionSequence implements CharSequence {
float posYInit;
float posYEnd;
if (textPositions.get(0).getRotation() == 90 && textPositions.get(0).getDir() != 0.0f) {
posXEnd = textPositions.get(0).getYDirAdj() + 2;
posYInit = getY1();
posYEnd = textPositions.get(textPositions.size() - 1).getXDirAdj() - height + 4;
if (textPositions.get(0).getRotation() == 0 && textPositions.get(0).getDir() == 90f) {
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 270.0f) {
posYInit = textPositions.get(0).getPageHeight() - getX1();
posYEnd = textPositions.get(0).getPageHeight() - getX2() - textPositions.get(0)
.getWidth() - textPositions.get(textPositions.size() - 1).getWidth() - 1;
posXInit = textPositions.get(0).getPageWidth() - textPositions.get(0).getYDirAdj() - 2;
posXEnd = textPositions.get(0).getPageWidth() - textPositions.get(textPositions.size() - 1)
.getYDirAdj() + height;
posYInit = getX1();
posYEnd = getX2() + textPositions.get(0).getWidthDirAdj() - textPositions.get(textPositions.size() - 1)
.getWidthDirAdj() - 3;
posXInit = textPositions.get(0).getYDirAdj() + 2;
posXEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height;
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 0.0f) {
posYInit = textPositions.get(0).getPageHeight() - textPositions.get(0).getYDirAdj() - 2;
posYEnd = posYInit + 1;
posXInit = textPositions.get(0).getXDirAdj();
posXEnd = textPositions.get(textPositions.size() - 1).getXDirAdj() + textPositions.get(textPositions.size() - 1).getWidthDirAdj() + 0.1f;
} else if (textPositions.get(0).getRotation() == 0 && textPositions.get(0).getDir() == 180f) {
posXInit = textPositions.get(0).getPageWidth() - getX1() + 1;
posXEnd = textPositions.get(0).getPageWidth() - getX2() + textPositions.get(0)
.getWidthDirAdj() - textPositions.get(textPositions.size() - 1).getWidthDirAdj() - 3;
posYInit = textPositions.get(0).getYDirAdj() - height + 2;
posYEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height + 2;
} else if (textPositions.get(0).getRotation() == 0 && textPositions.get(0).getDir() == 270f) {
posYInit = textPositions.get(0).getPageHeight() - getX1();
posYEnd = textPositions.get(0).getPageHeight() - getX2() - textPositions.get(0)
.getWidthDirAdj() - textPositions.get(textPositions.size() - 1).getWidthDirAdj() - 3;
@@ -292,6 +298,7 @@ public class TextPositionSequence implements CharSequence {
.getYDirAdj() + height;
} else if (textPositions.get(0).getRotation() == 90 && textPositions.get(0).getDir() == 0.0f) {
posXInit = textPositions.get(textPositions.size() - 1)
.getXDirAdj() + textPositions.get(textPositions.size() - 1).getHeightDir();
posXEnd = textPositions.get(0).getXDirAdj();
@@ -299,22 +306,93 @@ public class TextPositionSequence implements CharSequence {
posYEnd = textPositions.get(0).getPageHeight() - textPositions.get(textPositions.size() - 1)
.getYDirAdj() + 2;
} else if (textPositions.get(0).getRotation() == 0 && textPositions.get(0).getDir() == 90f) {
} else if (textPositions.get(0).getRotation() == 90 && textPositions.get(0).getDir() == 90.0f) {
posXEnd = textPositions.get(0).getYDirAdj() + 2;
posYInit = getY1();
posYEnd = textPositions.get(textPositions.size() - 1).getXDirAdj() - height + 4;
} else if (textPositions.get(0).getRotation() == 90 && textPositions.get(0).getDir() == 180.0f) {
posXInit = textPositions.get(0).getPageWidth() - textPositions.get(textPositions.size() - 1)
.getXDirAdj() - 4;
posXEnd = textPositions.get(0).getPageWidth() - textPositions.get(0).getXDirAdj();
posYInit = textPositions.get(0).getYDirAdj() - 2 - textPositions.get(textPositions.size() - 1)
.getHeightDir();
posYEnd = textPositions.get(textPositions.size() - 1)
.getYDirAdj() - textPositions.get(textPositions.size() - 1).getHeightDir();
} else if (textPositions.get(0).getRotation() == 90 && textPositions.get(0).getDir() == 270.0f) {
posXInit = textPositions.get(0).getPageWidth() - getX1();
posXEnd = textPositions.get(0).getPageWidth() - textPositions.get(0).getYDirAdj() - 2;
posYInit = textPositions.get(0).getPageHeight() - getY1();
posYEnd = textPositions.get(0).getPageHeight() - textPositions.get(textPositions.size() - 1)
.getXDirAdj() - height - 4;
} else if (textPositions.get(0).getRotation() == 180 && textPositions.get(0).getDir() == 0f) {
posXEnd = textPositions.get(textPositions.size() - 1)
.getXDirAdj() + textPositions.get(textPositions.size() - 1).getWidthDirAdj() + 1;
posYInit = textPositions.get(0).getPageHeight() - textPositions.get(0).getYDirAdj() - 2;
posYEnd = textPositions.get(0).getPageHeight() - textPositions.get(textPositions.size() - 1)
.getYDirAdj() + 2;
} else if (textPositions.get(0).getRotation() == 180 && textPositions.get(0).getDir() == 90f) {
posYInit = getX1();
posYEnd = getX2() + textPositions.get(0).getWidthDirAdj() - textPositions.get(textPositions.size() - 1)
.getWidthDirAdj() - 3;
posYEnd = getX2() - 3;
posXInit = textPositions.get(0).getYDirAdj() + 2;
posXEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height;
} else if (textPositions.get(0).getRotation() == 180 && textPositions.get(0).getDir() == 180f) {
posXInit = textPositions.get(0).getPageWidth() - getX1() + 1;
posXEnd = textPositions.get(0).getPageWidth() - getX2() + textPositions.get(0).getWidthDirAdj() - textPositions.get(textPositions.size() - 1)
.getWidthDirAdj() - 3;
posXEnd = textPositions.get(0).getPageWidth() - getX2() + textPositions.get(0)
.getWidthDirAdj() - textPositions.get(textPositions.size() - 1).getWidthDirAdj() - 3;
posYInit = textPositions.get(0).getYDirAdj() - height + 2;
posYEnd = textPositions.get(textPositions.size() - 1)
.getYDirAdj() - height + 2;
posYEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height + 2;
} else if (textPositions.get(0).getRotation() == 180 && textPositions.get(0).getDir() == 270.0f) {
posYInit = textPositions.get(0).getPageHeight() - getX1();
posYEnd = textPositions.get(0).getPageHeight() - getX2() - textPositions.get(0)
.getWidthDirAdj() - textPositions.get(textPositions.size() - 1).getWidthDirAdj();
posXInit = textPositions.get(0).getPageWidth() - textPositions.get(0).getYDirAdj() - 2;
posXEnd = textPositions.get(0).getPageWidth() - textPositions.get(textPositions.size() - 1)
.getYDirAdj() + height;
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 0.0f) {
posYInit = textPositions.get(0).getPageHeight() - textPositions.get(0).getYDirAdj() - 2;
posYEnd = posYInit + 1;
posXInit = textPositions.get(0).getXDirAdj();
posXEnd = textPositions.get(textPositions.size() - 1)
.getXDirAdj() + textPositions.get(textPositions.size() - 1).getWidthDirAdj() + 0.1f;
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 90.0f) {
posYInit = getX1();
posYEnd = getX2() - height;
posXInit = textPositions.get(0).getYDirAdj() + 2;
posXEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height;
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 180.0f) {
posXInit = textPositions.get(0).getPageWidth() - getX1() + 1;
posXEnd = textPositions.get(0).getPageWidth() - getX2() - 4;
posYInit = textPositions.get(0).getYDirAdj() - height + 2;
posYEnd = textPositions.get(textPositions.size() - 1).getYDirAdj() - height + 2;
} else if (textPositions.get(0).getRotation() == 270 && textPositions.get(0).getDir() == 270.0f) {
posYInit = textPositions.get(0).getPageHeight() - getX1();
posYEnd = textPositions.get(0).getPageHeight() - getX2() - height;
posXInit = textPositions.get(0).getPageWidth() - textPositions.get(0).getYDirAdj() - 2;
posXEnd = textPositions.get(0).getPageWidth() - textPositions.get(textPositions.size() - 1)
.getYDirAdj() + height;
} else {
// page rotation = 0 and text direction = 0
posXEnd = textPositions.get(textPositions.size() - 1)
.getXDirAdj() + textPositions.get(textPositions.size() - 1).getWidthDirAdj() + 1;
posYInit = textPositions.get(0).getPageHeight() - textPositions.get(0).getYDirAdj() - 2;
@@ -1,17 +1,23 @@
package com.iqser.red.service.redaction.v1.server.redaction.model;
import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.server.parsing.model.TextPositionSequence;
import lombok.Data;
import lombok.EqualsAndHashCode;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.server.parsing.model.TextPositionSequence;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.EqualsAndHashCode;
import lombok.NoArgsConstructor;
@Data
@Builder
@NoArgsConstructor
@AllArgsConstructor
@EqualsAndHashCode(onlyExplicitlyIncluded = true)
public class Entity implements ReasonHolder {
@@ -50,6 +56,8 @@ public class Entity implements ReasonHolder {
private EntityType entityType;
private boolean resized;
public Entity(String word, String type, boolean redaction, String redactionReason,
List<EntityPositionSequence> positionSequences, String headline, int matchedRule, int sectionNumber,
@@ -1,15 +1,18 @@
package com.iqser.red.service.redaction.v1.server.redaction.model;
import com.iqser.red.service.redaction.v1.server.parsing.model.TextPositionSequence;
import lombok.AllArgsConstructor;
import lombok.Data;
import lombok.EqualsAndHashCode;
import lombok.RequiredArgsConstructor;
import java.util.ArrayList;
import java.util.List;
import com.iqser.red.service.redaction.v1.server.parsing.model.TextPositionSequence;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.EqualsAndHashCode;
import lombok.RequiredArgsConstructor;
@Data
@Builder
@RequiredArgsConstructor
@AllArgsConstructor
@EqualsAndHashCode
@@ -13,6 +13,7 @@ public class PdfImage {
@NonNull
private ImageType imageType;
private boolean isAppendedToParagraph;
@NonNull
private boolean hasTransparency;
@NonNull
private int page;
@@ -1,14 +1,16 @@
package com.iqser.red.service.redaction.v1.server.redaction.model;
import com.dslplatform.json.CompiledJson;
import com.dslplatform.json.JsonAttribute;
import com.fasterxml.jackson.annotation.JsonIgnore;
import lombok.AllArgsConstructor;
import lombok.Builder;
import lombok.Data;
import lombok.NoArgsConstructor;
@Data
@Builder
@CompiledJson
@NoArgsConstructor
@AllArgsConstructor
@@ -1,28 +1,38 @@
package com.iqser.red.service.redaction.v1.server.redaction.model;
import com.iqser.red.service.redaction.v1.model.ArgumentType;
import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.model.FileAttribute;
import com.iqser.red.service.redaction.v1.server.classification.model.TextBlock;
import com.iqser.red.service.redaction.v1.server.redaction.utils.EntitySearchUtils;
import com.iqser.red.service.redaction.v1.server.redaction.utils.FindEntityDetails;
import com.iqser.red.service.redaction.v1.server.redaction.utils.Patterns;
import com.iqser.red.service.redaction.v1.server.redaction.utils.SearchImplementation;
import lombok.Builder;
import lombok.Data;
import lombok.extern.slf4j.Slf4j;
import org.apache.commons.lang3.StringUtils;
import java.lang.annotation.ElementType;
import java.lang.annotation.Retention;
import java.lang.annotation.RetentionPolicy;
import java.lang.annotation.Target;
import java.util.*;
import java.util.ArrayList;
import java.util.Collection;
import java.util.Comparator;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.regex.Matcher;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
import org.apache.commons.lang3.StringUtils;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.ManualRedactions;
import com.iqser.red.service.redaction.v1.model.ArgumentType;
import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.model.FileAttribute;
import com.iqser.red.service.redaction.v1.model.SectionArea;
import com.iqser.red.service.redaction.v1.server.classification.model.TextBlock;
import com.iqser.red.service.redaction.v1.server.redaction.utils.EntitySearchUtils;
import com.iqser.red.service.redaction.v1.server.redaction.utils.FindEntityDetails;
import com.iqser.red.service.redaction.v1.server.redaction.utils.Patterns;
import com.iqser.red.service.redaction.v1.server.redaction.utils.SearchImplementation;
import lombok.Builder;
import lombok.Data;
import lombok.extern.slf4j.Slf4j;
@Data
@Slf4j
@Builder
@@ -61,15 +71,22 @@ public class Section {
@Builder.Default
private List<FileAttribute> fileAttributes = new ArrayList<>();
@Builder.Default
private List<SectionArea> sectionAreas = new ArrayList<>();
private ManualRedactions manualRedactions;
@SuppressWarnings("unused")
@WhenCondition
public void addAiEntities(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.TYPE) String asType) {
Set<Entity> entitiesOfType = nerEntities.stream().filter(nerEntity -> nerEntity.getType().equals(type)).collect(Collectors.toSet());
Set<Entity> entitiesOfType = nerEntities.stream()
.filter(nerEntity -> nerEntity.getType().equals(type))
.collect(Collectors.toSet());
List<String> values = entitiesOfType.stream().map(Entity::getWord).collect(Collectors.toList());
Set<Entity> found = EntitySearchUtils.findEntities(searchText, new SearchImplementation(values, dictionary.isCaseInsensitiveDictionary(asType)), dictionary.getType(asType), new FindEntityDetails(asType, headline, sectionNumber, false, false, Engine.NER, EntityType.RECOMMENDATION));
EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary);
EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary, manualRedactions);
Set<Entity> finalResult = new HashSet<>();
@@ -93,15 +110,21 @@ public class Section {
nerEntities.removeAll(entitiesOfType);
}
@SuppressWarnings("unused")
@WhenCondition
public void combineAiTypes(@Argument(ArgumentType.TYPE) String startType, @Argument(ArgumentType.TYPE) String combineTypes,
@Argument(ArgumentType.INTEGER) int maxDistanceBetween, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.INTEGER) int minPartMatches, @Argument(ArgumentType.BOOLEAN) boolean allowDuplicateTypes) {
public void combineAiTypes(@Argument(ArgumentType.TYPE) String startType,
@Argument(ArgumentType.TYPE) String combineTypes,
@Argument(ArgumentType.INTEGER) int maxDistanceBetween,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.INTEGER) int minPartMatches,
@Argument(ArgumentType.BOOLEAN) boolean allowDuplicateTypes) {
Set<String> combineSet = Set.of(combineTypes.split(","));
List<Entity> sorted = nerEntities.stream().sorted(Comparator.comparing(Entity::getStart)).collect(Collectors.toList());
List<Entity> sorted = nerEntities.stream()
.sorted(Comparator.comparing(Entity::getStart))
.collect(Collectors.toList());
Set<Entity> found = new HashSet<>();
int start = -1;
int lastEnd = -1;
@@ -158,49 +181,67 @@ public class Section {
}
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByIdEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String id, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByIdEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String id,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream().anyMatch(attribute -> id.equals(attribute.getId()) && value.equals(attribute.getValue()));
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> id.equals(attribute.getId()) && value.equals(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByPlaceholderEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String placeholder, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByPlaceholderEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String placeholder,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream().anyMatch(attribute -> placeholder.equals(attribute.getPlaceholder()) && value.equals(attribute.getValue()));
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> placeholder.equals(attribute.getPlaceholder()) && value.equals(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByLabelEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String label, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByLabelEquals(@Argument(ArgumentType.FILE_ATTRIBUTE) String label,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream().anyMatch(attribute -> label.equals(attribute.getLabel()) && value.equals(attribute.getValue()));
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> label.equals(attribute.getLabel()) && value.equals(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByIdEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String id, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByIdEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String id,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream().anyMatch(attribute -> id.equals(attribute.getId()) && value.equalsIgnoreCase(attribute.getValue()));
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> id.equals(attribute.getId()) && value.equalsIgnoreCase(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByPlaceholderEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String placeholder, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByPlaceholderEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String placeholder,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> placeholder.equals(attribute.getPlaceholder()) && value.equalsIgnoreCase(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean fileAttributeByLabelEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String label, @Argument(ArgumentType.STRING) String value) {
public boolean fileAttributeByLabelEqualsIgnoreCase(@Argument(ArgumentType.FILE_ATTRIBUTE) String label,
@Argument(ArgumentType.STRING) String value) {
return fileAttributes != null && fileAttributes.stream().anyMatch(attribute -> label.equals(attribute.getLabel()) && value.equalsIgnoreCase(attribute.getValue()));
return fileAttributes != null && fileAttributes.stream()
.anyMatch(attribute -> label.equals(attribute.getLabel()) && value.equalsIgnoreCase(attribute.getValue()));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean hasTableHeader(@Argument(ArgumentType.STRING) String headerName) {
@@ -209,6 +250,7 @@ public class Section {
return tabularData != null && tabularData.containsKey(cleanHeaderName);
}
@SuppressWarnings("unused")
@WhenCondition
public boolean aiMatchesType(@Argument(ArgumentType.TYPE) String type) {
@@ -216,6 +258,7 @@ public class Section {
return nerEntities.stream().anyMatch(entity -> !entity.isIgnored() && entity.getType().equals(type));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean matchesType(@Argument(ArgumentType.TYPE) String type) {
@@ -223,6 +266,7 @@ public class Section {
return entities.stream().anyMatch(entity -> !entity.isIgnored() && entity.getType().equals(type));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean matchesImageType(@Argument(ArgumentType.TYPE) String type) {
@@ -230,6 +274,7 @@ public class Section {
return images.stream().anyMatch(image -> !image.isIgnored() && image.getType().equals(type));
}
@SuppressWarnings("unused")
@WhenCondition
public boolean headlineContainsWord(@Argument(ArgumentType.STRING) String word) {
@@ -237,9 +282,11 @@ public class Section {
return StringUtils.containsIgnoreCase(headline, word);
}
@SuppressWarnings("unused")
@WhenCondition
public boolean containsRegEx(@Argument(ArgumentType.STRING) String regEx, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive) {
public boolean containsRegEx(@Argument(ArgumentType.STRING) String regEx,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive) {
var compiledPattern = Patterns.getCompiledPattern(regEx, patternCaseInsensitive);
@@ -248,19 +295,26 @@ public class Section {
return matcher.find();
}
@SuppressWarnings("unused")
@WhenCondition
public boolean rowEquals(@Argument(ArgumentType.STRING) String headerName, @Argument(ArgumentType.STRING) String value) {
public boolean rowEquals(@Argument(ArgumentType.STRING) String headerName,
@Argument(ArgumentType.STRING) String value) {
String cleanHeaderName = headerName.replaceAll("\n", "").replaceAll(" ", "").replaceAll("-", "");
return tabularData != null && tabularData.containsKey(cleanHeaderName) && tabularData.get(cleanHeaderName).toString().equals(value);
return tabularData != null && tabularData.containsKey(cleanHeaderName) && tabularData.get(cleanHeaderName)
.toString()
.equals(value);
}
@SuppressWarnings("unused")
@ThenAction
public void expandByPrefixRegEx(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.REGEX) String prefixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive, @Argument(ArgumentType.INTEGER) int group) {
@SuppressWarnings("unused")
public void expandByPrefixRegEx(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.REGEX) String prefixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group) {
expandByPrefixRegEx(type, prefixPattern, patternCaseInsensitive, group, null);
}
@@ -268,11 +322,15 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void expandByPrefixRegEx(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.REGEX) String prefixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive, @Argument(ArgumentType.INTEGER) int group,
public void expandByPrefixRegEx(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.REGEX) String prefixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.REGEX) String valuePattern) {
if (StringUtils.isEmpty(prefixPattern)) return;
if (StringUtils.isEmpty(prefixPattern)) {
return;
}
var compiledValuePattern = valuePattern == null ? null : Patterns.getCompiledPattern(valuePattern, patternCaseInsensitive);
var compiledPrefixPattern = Patterns.getCompiledPattern(prefixPattern, patternCaseInsensitive);
@@ -317,8 +375,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void expandByRegEx(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.REGEX) String suffixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive, @Argument(ArgumentType.INTEGER) int group) {
public void expandByRegEx(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.REGEX) String suffixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group) {
expandByRegEx(type, suffixPattern, patternCaseInsensitive, group, null);
}
@@ -326,11 +386,15 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void expandByRegEx(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.REGEX) String suffixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive, @Argument(ArgumentType.INTEGER) int group,
public void expandByRegEx(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.REGEX) String suffixPattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.REGEX) String valuePattern) {
if (StringUtils.isEmpty(suffixPattern)) return;
if (StringUtils.isEmpty(suffixPattern)) {
return;
}
var compiledValuePattern = valuePattern == null ? null : Patterns.getCompiledPattern(valuePattern, patternCaseInsensitive);
var compiledSuffixPattern = Patterns.getCompiledPattern(suffixPattern, patternCaseInsensitive);
@@ -374,7 +438,9 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactImage(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason,
public void redactImage(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactImage(type, ruleNumber, reason, legalBasis, true);
@@ -383,7 +449,9 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotImage(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason) {
public void redactNotImage(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason) {
redactImage(type, ruleNumber, reason, null, false);
}
@@ -391,7 +459,8 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redact(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason,
public void redact(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redact(type, ruleNumber, reason, legalBasis, true);
@@ -400,7 +469,8 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNot(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason) {
public void redactNot(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason) {
redact(type, ruleNumber, reason, null, false);
}
@@ -408,8 +478,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactLineAfter(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere, @Argument(ArgumentType.STRING) String reason,
public void redactLineAfter(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactLineAfter(start, asType, ruleNumber, redactEverywhere, reason, legalBasis, true);
@@ -418,8 +490,11 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotLineAfter(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere, @Argument(ArgumentType.STRING) String reason) {
public void redactNotLineAfter(@Argument(ArgumentType.STRING) String start,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason) {
redactLineAfter(start, asType, ruleNumber, redactEverywhere, reason, null, false);
@@ -428,9 +503,12 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void redactByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactByRegEx(pattern, patternCaseInsensitive, group, asType, ruleNumber, reason, legalBasis, true);
}
@@ -438,8 +516,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
public void redactNotByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason) {
redactByRegEx(pattern, patternCaseInsensitive, group, asType, ruleNumber, reason, null, false);
@@ -448,9 +528,12 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactBetween(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.STRING) String stop, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void redactBetween(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.STRING) String stop,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactBetween(start, stop, asType, ruleNumber, redactEverywhere, reason, legalBasis, true);
}
@@ -458,8 +541,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotBetween(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.STRING) String stop, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
public void redactNotBetween(@Argument(ArgumentType.STRING) String start,
@Argument(ArgumentType.STRING) String stop, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason) {
redactBetween(start, stop, asType, ruleNumber, redactEverywhere, reason, null, false);
@@ -468,9 +553,13 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactLinesBetween(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.STRING) String stop, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void redactLinesBetween(@Argument(ArgumentType.STRING) String start,
@Argument(ArgumentType.STRING) String stop,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactLinesBetween(start, stop, asType, ruleNumber, redactEverywhere, reason, legalBasis, true);
}
@@ -478,8 +567,11 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotLinesBetween(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.STRING) String stop, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
public void redactNotLinesBetween(@Argument(ArgumentType.STRING) String start,
@Argument(ArgumentType.STRING) String stop,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.BOOLEAN) boolean redactEverywhere,
@Argument(ArgumentType.STRING) String reason) {
redactLinesBetween(start, stop, asType, ruleNumber, redactEverywhere, reason, null, false);
@@ -488,8 +580,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactCell(@Argument(ArgumentType.STRING) String cellHeader, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.BOOLEAN) boolean addAsRecommendations, @Argument(ArgumentType.STRING) String reason,
public void redactCell(@Argument(ArgumentType.STRING) String cellHeader,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.BOOLEAN) boolean addAsRecommendations,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
annotateCell(cellHeader, ruleNumber, type, true, addAsRecommendations, reason, legalBasis);
@@ -498,8 +592,11 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotCell(@Argument(ArgumentType.STRING) String cellHeader, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.BOOLEAN) boolean addAsRecommendations, @Argument(ArgumentType.STRING) String reason) {
public void redactNotCell(@Argument(ArgumentType.STRING) String cellHeader,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.BOOLEAN) boolean addAsRecommendations,
@Argument(ArgumentType.STRING) String reason) {
annotateCell(cellHeader, ruleNumber, type, false, addAsRecommendations, reason, null);
}
@@ -507,9 +604,13 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactAndRecommendByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void redactAndRecommendByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
redactAndRecommendByRegEx(pattern, patternCaseInsensitive, group, asType, ruleNumber, reason, legalBasis, true);
}
@@ -517,9 +618,12 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotAndRecommendByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason) {
public void redactNotAndRecommendByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason) {
redactAndRecommendByRegEx(pattern, patternCaseInsensitive, group, asType, ruleNumber, reason, null, false);
}
@@ -527,8 +631,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void addRecommendationByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType) {
public void addRecommendationByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.TYPE) String asType) {
Pattern compiledPattern = Patterns.getCompiledPattern(pattern, patternCaseInsensitive);
@@ -545,11 +651,14 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactNotAndReference(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.REFERENCE_TYPE) String referenceType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.STRING) String reason) {
public void redactNotAndReference(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.REFERENCE_TYPE) String referenceType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason) {
Set<Entity> references = entities.stream().filter(entity -> entity.getType().equals(referenceType)).collect(Collectors.toSet());
Set<Entity> references = entities.stream()
.filter(entity -> entity.getType().equals(referenceType))
.collect(Collectors.toSet());
entities.forEach(entity -> {
if (entity.getType().equals(type)) {
@@ -564,8 +673,11 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void redactIfPrecededBy(@Argument(ArgumentType.STRING) String prefix, @Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void redactIfPrecededBy(@Argument(ArgumentType.STRING) String prefix,
@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
entities.forEach(entity -> {
if (entity.getType().equals(type) && searchText.indexOf(prefix + entity.getWord()) != 1) {
@@ -580,13 +692,16 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void addRedaction(@Argument(ArgumentType.STRING) String value, @Argument(ArgumentType.TYPE) String asType, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason, @Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
public void addRedaction(@Argument(ArgumentType.STRING) String value, @Argument(ArgumentType.TYPE) String asType,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
Set<Entity> found = findEntities(value.trim(), asType, true, true, ruleNumber, reason, legalBasis, Engine.RULE, false);
EntitySearchUtils.addEntitiesIgnoreRank(entities, found);
}
@ThenAction
@SuppressWarnings("unused")
public void ignore(@Argument(ArgumentType.TYPE) String type) {
@@ -594,18 +709,22 @@ public class Section {
entities.removeIf(entity -> entity.getType().equals(type) && entity.getEntityType().equals(EntityType.ENTITY));
}
@ThenAction
@SuppressWarnings("unused")
public void ignoreRecommendations(@Argument(ArgumentType.TYPE) String type) {
entities.removeIf(entity -> entity.getType().equals(type) && entity.getEntityType().equals(EntityType.RECOMMENDATION));
entities.removeIf(entity -> entity.getType().equals(type) && entity.getEntityType()
.equals(EntityType.RECOMMENDATION));
}
@ThenAction
@SuppressWarnings("unused")
public void expandToFalsePositiveByRegEx(@Argument(ArgumentType.TYPE) String type, @Argument(ArgumentType.STRING) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive, @Argument(ArgumentType.INTEGER) int group) {
public void expandToFalsePositiveByRegEx(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.STRING) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group) {
Pattern compiledPattern = Patterns.getCompiledPattern(pattern, patternCaseInsensitive);
@@ -634,8 +753,10 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void addHintAnnotationByRegEx(@Argument(ArgumentType.REGEX) String pattern, @Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group, @Argument(ArgumentType.TYPE) String asType) {
public void addHintAnnotationByRegEx(@Argument(ArgumentType.REGEX) String pattern,
@Argument(ArgumentType.BOOLEAN) boolean patternCaseInsensitive,
@Argument(ArgumentType.INTEGER) int group,
@Argument(ArgumentType.TYPE) String asType) {
Pattern compiledPattern = Patterns.getCompiledPattern(pattern, patternCaseInsensitive);
@@ -653,7 +774,8 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void addHintAnnotation(@Argument(ArgumentType.STRING) String value, @Argument(ArgumentType.TYPE) String asType) {
public void addHintAnnotation(@Argument(ArgumentType.STRING) String value,
@Argument(ArgumentType.TYPE) String asType) {
Set<Entity> found = findEntities(value.trim(), asType, true, false, 0, null, null, Engine.RULE, false);
EntitySearchUtils.addEntitiesIgnoreRank(entities, found);
@@ -662,7 +784,8 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void recommendLineAfter(@Argument(ArgumentType.STRING) String start, @Argument(ArgumentType.TYPE) String asType) {
public void recommendLineAfter(@Argument(ArgumentType.STRING) String start,
@Argument(ArgumentType.TYPE) String asType) {
String[] values = StringUtils.substringsBetween(text, start, "\n");
@@ -687,14 +810,53 @@ public class Section {
@ThenAction
@SuppressWarnings("unused")
public void highlightCell(@Argument(ArgumentType.STRING) String cellHeader, @Argument(ArgumentType.RULE_NUMBER) int ruleNumber, @Argument(ArgumentType.TYPE) String type) {
public void highlightCell(@Argument(ArgumentType.STRING) String cellHeader,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.TYPE) String type) {
annotateCell(cellHeader, ruleNumber, type, false, false, null, null);
}
private void redactAndRecommendByRegEx(String pattern, boolean patternCaseInsensitive, int group, String asType, int ruleNumber, String reason, String legalBasis,
boolean redaction) {
@ThenAction
@SuppressWarnings("unused")
public void redactSection(@Argument(ArgumentType.TYPE) String type,
@Argument(ArgumentType.RULE_NUMBER) int ruleNumber,
@Argument(ArgumentType.STRING) String reason,
@Argument(ArgumentType.LEGAL_BASIS) String legalBasis) {
for (SectionArea sectionArea : sectionAreas) {
RedRectangle2D position = RedRectangle2D.builder()
.height(sectionArea.getHeight() + 4)
.width(sectionArea.getWidth() + 4)
.x(sectionArea.getTopLeft().getX() - 2)
.y(sectionArea.getTopLeft().getY() - 2)
.build();
log.debug("SectionArea: {}", sectionArea);
log.debug("Position {}", position.toString());
Image image = Image.builder()
.page(sectionArea.getPage())
.position(position)
.redaction(true)
.hasTransparency(false)
.sectionNumber(sectionNumber)
.section(headline)
.matchedRule(ruleNumber)
.legalBasis(legalBasis)
.redactionReason(reason)
.type(type)
.build();
images.add(image);
}
}
private void redactAndRecommendByRegEx(String pattern, boolean patternCaseInsensitive, int group, String asType,
int ruleNumber, String reason, String legalBasis, boolean redaction) {
Pattern compiledPattern = Patterns.getCompiledPattern(pattern, patternCaseInsensitive);
Matcher matcher = compiledPattern.matcher(searchText);
@@ -709,11 +871,12 @@ public class Section {
}
private Set<Entity> findEntities(String value, String asType, boolean caseInsensitive, boolean redacted, int ruleNumber, String reason, String legalBasis, Engine engine, boolean asRecommendation) {
private Set<Entity> findEntities(String value, String asType, boolean caseInsensitive, boolean redacted,
int ruleNumber, String reason, String legalBasis, Engine engine,
boolean asRecommendation) {
String text = caseInsensitive ? searchText.toLowerCase() : searchText;
Set<Entity> found = EntitySearchUtils.findEntities(text, new SearchImplementation(value, caseInsensitive), dictionary.getType(asType),
new FindEntityDetails(asType, headline, sectionNumber, false, false, engine, asRecommendation ? EntityType.RECOMMENDATION : EntityType.ENTITY));
Set<Entity> found = EntitySearchUtils.findEntities(text, new SearchImplementation(value, caseInsensitive), dictionary.getType(asType), new FindEntityDetails(asType, headline, sectionNumber, false, false, engine, asRecommendation ? EntityType.RECOMMENDATION : EntityType.ENTITY));
found.forEach(entity -> {
if (redacted) {
entity.setRedaction(true);
@@ -723,10 +886,13 @@ public class Section {
}
});
return EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary);
return EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary, manualRedactions);
}
private void redact(String type, int ruleNumber, String reason, String legalBasis, boolean redaction) {
entities.forEach(entity -> {
@@ -753,7 +919,8 @@ public class Section {
}
private void annotateCell(String cellHeader, int ruleNumber, String type, boolean redact, boolean addAsRecommendations, String reason, String legalBasis) {
private void annotateCell(String cellHeader, int ruleNumber, String type, boolean redact,
boolean addAsRecommendations, String reason, String legalBasis) {
String cleanHeaderName = cellHeader.replaceAll("\n", "").replaceAll(" ", "").replaceAll("-", "");
@@ -776,7 +943,8 @@ public class Section {
Set<Entity> singleEntitySet = new HashSet<>();
singleEntitySet.add(entity);
EntitySearchUtils.clearAndFindPositions(singleEntitySet, searchableText, dictionary);
EntitySearchUtils.clearAndFindPositions(singleEntitySet, searchableText, dictionary, manualRedactions);
EntitySearchUtils.addEntitiesWithHigherRank(entities, entity, dictionary);
@@ -800,7 +968,8 @@ public class Section {
}
private void redactLineAfter(String start, String asType, int ruleNumber, boolean redactEverywhere, String reason, String legalBasis, boolean redaction) {
private void redactLineAfter(String start, String asType, int ruleNumber, boolean redactEverywhere, String reason,
String legalBasis, boolean redaction) {
String[] values = StringUtils.substringsBetween(text, start, "\n");
@@ -819,7 +988,8 @@ public class Section {
}
private void redactByRegEx(String pattern, boolean patternCaseInsensitive, int group, String asType, int ruleNumber, String reason, String legalBasis, boolean redaction) {
private void redactByRegEx(String pattern, boolean patternCaseInsensitive, int group, String asType, int ruleNumber,
String reason, String legalBasis, boolean redaction) {
Pattern compiledPattern = Patterns.getCompiledPattern(pattern, patternCaseInsensitive);
@@ -835,7 +1005,8 @@ public class Section {
}
private void redactBetween(String start, String stop, String asType, int ruleNumber, boolean redactEverywhere, String reason, String legalBasis, boolean redaction) {
private void redactBetween(String start, String stop, String asType, int ruleNumber, boolean redactEverywhere,
String reason, String legalBasis, boolean redaction) {
String[] values = StringUtils.substringsBetween(searchText, start, stop);
@@ -855,7 +1026,8 @@ public class Section {
}
private void redactLinesBetween(String start, String stop, String asType, int ruleNumber, boolean redactEverywhere, String reason, String legalBasis, boolean redaction) {
private void redactLinesBetween(String start, String stop, String asType, int ruleNumber, boolean redactEverywhere,
String reason, String legalBasis, boolean redaction) {
String[] values = StringUtils.substringsBetween(text, start, stop);
@@ -1,5 +1,7 @@
package com.iqser.red.service.redaction.v1.server.redaction.model;
import java.util.List;
import lombok.AllArgsConstructor;
import lombok.Data;
@@ -9,5 +11,6 @@ public class SectionSearchableTextPair {
private Section section;
private SearchableText searchableText;
private List<Integer> cellStarts;
}
@@ -11,4 +11,5 @@ public class ImageMetadata {
private Position position;
private Geometry geometry;
private Filters filters;
private boolean alpha;
}
@@ -5,6 +5,7 @@ import com.iqser.red.service.persistence.service.v1.api.model.annotations.entity
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualForceRedaction;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualImageRecategorization;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualLegalBasisChange;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualResizeRedaction;
import com.iqser.red.service.persistence.service.v1.api.model.dossiertemplate.dossier.file.FileType;
import com.iqser.red.service.persistence.service.v1.api.model.dossiertemplate.legalbasis.LegalBasis;
import com.iqser.red.service.redaction.v1.model.*;
@@ -57,7 +58,7 @@ public class AnalyzeService {
private final ImportedRedactionService importedRedactionService;
@Timed("analyzeDocumentStructure")
@Timed("redactmanager_analyzeDocumentStructure")
public AnalyzeResult analyzeDocumentStructure(StructureAnalyzeRequest analyzeRequest) {
long startTime = System.currentTimeMillis();
@@ -106,7 +107,7 @@ public class AnalyzeService {
}
@Timed("reanalyze")
@Timed("redactmanager_reanalyze")
@SneakyThrows
public AnalyzeResult reanalyze(@RequestBody AnalyzeRequest analyzeRequest) {
@@ -120,14 +121,10 @@ public class AnalyzeService {
return analyze(analyzeRequest);
}
// var dis = System.currentTimeMillis();
DictionaryIncrement dictionaryIncrement = dictionaryService.getDictionaryIncrements(analyzeRequest.getDossierTemplateId(), new DictionaryVersion(redactionLog.getDictionaryVersion(), redactionLog.getDossierDictionaryVersion()), analyzeRequest.getDossierId());
// log.info("Dictionary Increment time time: {} ms", (System.currentTimeMillis() - dis));
// var fis = System.currentTimeMillis();
Set<Integer> sectionsToReanalyse = !analyzeRequest.getSectionsToReanalyse()
.isEmpty() ? analyzeRequest.getSectionsToReanalyse() : findSectionsToReanalyse(dictionaryIncrement, redactionLog, text, analyzeRequest);
// log.info("Find sections time: {} ms", (System.currentTimeMillis() - fis));
if (sectionsToReanalyse.isEmpty()) {
return finalizeAnalysis(analyzeRequest, startTime, redactionLog, text, dictionaryIncrement.getDictionaryVersion(), true);
@@ -145,36 +142,26 @@ public class AnalyzeService {
.filter(sectionText -> sectionsToReanalyse.contains(sectionText.getSectionNumber()))
.collect(Collectors.toList());
// long kis = System.currentTimeMillis();
KieContainer kieContainer = droolsExecutionService.updateRules(analyzeRequest.getDossierTemplateId());
// log.info("Kie time: {} ms", (System.currentTimeMillis() - kis));
// long dds = System.currentTimeMillis();
Dictionary dictionary = dictionaryService.getDeepCopyDictionary(analyzeRequest.getDossierTemplateId(), analyzeRequest.getDossierId());
// log.info("Dict Time time: {} ms", (System.currentTimeMillis() - dds));
// long pis = System.currentTimeMillis();
PageEntities pageEntities = entityRedactionService.findEntities(dictionary, reanalysisSections, kieContainer, analyzeRequest, nerEntities);
// log.info("Find Entities time: {}", (System.currentTimeMillis() - pis));
// long crs = System.currentTimeMillis();
var newRedactionLogEntries = redactionLogCreatorService.createRedactionLog(pageEntities, text.getNumberOfPages(), analyzeRequest.getDossierTemplateId());
// log.info("Create Redaction-log time: {} ms", (System.currentTimeMillis() - crs));
// long prs = System.currentTimeMillis();
var importedRedactionFilteredEntries = importedRedactionService.processImportedRedactions(analyzeRequest.getDossierTemplateId(), analyzeRequest.getDossierId(), analyzeRequest.getFileId(), newRedactionLogEntries, false);
// log.info("Process imports time: {} ms", (System.currentTimeMillis() - prs));
redactionLog.getRedactionLogEntry()
.removeIf(entry -> sectionsToReanalyse.contains(entry.getSectionNumber()) && !entry.getType()
.equals(IMPORTED_REDACTION_TYPE));
redactionLog.getRedactionLogEntry().addAll(importedRedactionFilteredEntries);
// var fls = System.currentTimeMillis();
// log.info("Finalize time: {} ms", (System.currentTimeMillis() - fls));
return finalizeAnalysis(analyzeRequest, startTime, redactionLog, text, dictionaryIncrement.getDictionaryVersion(), true);
}
@Timed("analyze")
@Timed("redactmanager_analyze")
public AnalyzeResult analyze(AnalyzeRequest analyzeRequest) {
long startTime = System.currentTimeMillis();
@@ -207,6 +194,7 @@ public class AnalyzeService {
}
@Timed("redactmanager_findSectionsToReanalyse")
private Set<Integer> findSectionsToReanalyse(DictionaryIncrement dictionaryIncrement, RedactionLog redactionLog,
Text text, AnalyzeRequest analyzeRequest) {
@@ -224,7 +212,6 @@ public class AnalyzeService {
}
}
// long ss = System.currentTimeMillis();
var dictionaryIncrementsSearch = new SearchImplementation(dictionaryIncrement.getValues().stream()
.map(DictionaryIncrementValue::getValue).collect(Collectors.toList()), true);
@@ -236,7 +223,6 @@ public class AnalyzeService {
}
}
// log.info("Section Find time: {}", (System.currentTimeMillis() - ss));
log.info("Should reanalyze {} sections for request: {}, took: {}", sectionsToReanalyse.size(), analyzeRequest, System.currentTimeMillis() - start);
@@ -282,7 +268,9 @@ public class AnalyzeService {
return new HashSet<>();
}
return Stream.concat(manualRedactions.getLegalBasisChanges()
return Stream.concat(manualRedactions.getResizeRedactions()
.stream()
.map(ManualResizeRedaction::getAnnotationId), Stream.concat(manualRedactions.getLegalBasisChanges()
.stream()
.map(ManualLegalBasisChange::getAnnotationId), Stream.concat(manualRedactions.getImageRecategorization()
.stream()
@@ -290,7 +278,7 @@ public class AnalyzeService {
.stream()
.map(IdRemoval::getAnnotationId), manualRedactions.getForceRedactions()
.stream()
.map(ManualForceRedaction::getAnnotationId)))).collect(Collectors.toSet());
.map(ManualForceRedaction::getAnnotationId))))).collect(Collectors.toSet());
}
public List<RedactionLogLegalBasis> convert(List<LegalBasis> legalBasis) {
@@ -5,9 +5,12 @@ import com.iqser.red.service.persistence.service.v1.api.model.dossiertemplate.ty
import com.iqser.red.service.redaction.v1.server.client.DictionaryClient;
import com.iqser.red.service.redaction.v1.server.redaction.model.Dictionary;
import com.iqser.red.service.redaction.v1.server.redaction.model.*;
import feign.FeignException;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.apache.commons.collections4.CollectionUtils;
import org.apache.commons.lang3.SerializationUtils;
import org.springframework.stereotype.Service;
@@ -17,7 +20,6 @@ import java.util.List;
import java.util.*;
import java.util.stream.Collectors;
@Slf4j
@Service
@RequiredArgsConstructor
@@ -29,6 +31,7 @@ public class DictionaryService {
private final Map<String, DictionaryRepresentation> dictionariesByDossier = new HashMap<>();
@Timed("redactmanager_updateDictionary")
public DictionaryVersion updateDictionary(String dossierTemplateId, String dossierId) {
log.info("Updating dictionary data for dossierTemplate {} and dossier {}", dossierTemplateId, dossierId);
@@ -44,11 +47,16 @@ public class DictionaryService {
updateDictionaryEntry(dossierTemplateId, dossierDictionaryVersion, getVersion(dossierDictionary), dossierId);
}
return DictionaryVersion.builder().dossierTemplateVersion(dossierTemplateDictionaryVersion).dossierVersion(dossierDictionaryVersion).build();
return DictionaryVersion.builder()
.dossierTemplateVersion(dossierTemplateDictionaryVersion)
.dossierVersion(dossierDictionaryVersion)
.build();
}
public DictionaryIncrement getDictionaryIncrements(String dossierTemplateId, DictionaryVersion fromVersion, String dossierId) {
@Timed("redactmanager_getDictionaryIncrements")
public DictionaryIncrement getDictionaryIncrements(String dossierTemplateId, DictionaryVersion fromVersion,
String dossierId) {
DictionaryVersion version = updateDictionary(dossierTemplateId, dossierId);
@@ -102,58 +110,69 @@ public class DictionaryService {
try {
DictionaryRepresentation dictionaryRepresentation = new DictionaryRepresentation();
var typeResponse = dossierId == null ? dictionaryClient.getAllTypesForDossierTemplate(dossierTemplateId, false)
: dictionaryClient.getAllTypesForDossier(dossierId, false);
var typeResponse = dossierId == null ? dictionaryClient.getAllTypesForDossierTemplate(dossierTemplateId, false) : dictionaryClient.getAllTypesForDossier(dossierId, false);
if (CollectionUtils.isNotEmpty(typeResponse)) {
List<DictionaryModel> dictionary = typeResponse
.stream()
.map(t -> {
List<DictionaryModel> dictionary = typeResponse.stream().map(t -> {
Optional<DictionaryModel> oldModel;
if (dossierId == null) {
var representation = dictionariesByDossierTemplate.get(dossierTemplateId);
oldModel = representation != null ? representation.getDictionary().stream().filter(f -> f.getType().equals(t.getType())).findAny() : Optional.empty();
} else {
var representation = dictionariesByDossier.get(dossierId);
oldModel = representation != null ? representation.getDictionary().stream().filter(f -> f.getType().equals(t.getType())).findAny() : Optional.empty();
}
Optional<DictionaryModel> oldModel;
if (dossierId == null) {
var representation = dictionariesByDossierTemplate.get(dossierTemplateId);
oldModel = representation != null ? representation.getDictionary()
.stream()
.filter(f -> f.getType().equals(t.getType()))
.findAny() : Optional.empty();
} else {
var representation = dictionariesByDossier.get(dossierId);
oldModel = representation != null ? representation.getDictionary()
.stream()
.filter(f -> f.getType().equals(t.getType()))
.findAny() : Optional.empty();
}
Set<DictionaryEntry> entries = new HashSet<>();
Set<DictionaryEntry> falsePositives = new HashSet<>();
Set<DictionaryEntry> falseRecommendations = new HashSet<>();
Set<DictionaryEntry> entries = new HashSet<>();
Set<DictionaryEntry> falsePositives = new HashSet<>();
Set<DictionaryEntry> falseRecommendations = new HashSet<>();
DictionaryEntries newEntries = getEntries(t.getId(), currentVersion);
DictionaryEntries newEntries = getEntries(t.getId(), currentVersion);
var newValues = newEntries.getEntries().stream().map(DictionaryEntry::getValue).collect(Collectors.toSet());
var newFalsePositivesValues = newEntries.getFalsePositives().stream().map(DictionaryEntry::getValue).collect(Collectors.toSet());
var newFalseRecommendationsValues = newEntries.getFalseRecommendations().stream().map(DictionaryEntry::getValue).collect(Collectors.toSet());
var newValues = newEntries.getEntries()
.stream()
.map(DictionaryEntry::getValue)
.collect(Collectors.toSet());
var newFalsePositivesValues = newEntries.getFalsePositives()
.stream()
.map(DictionaryEntry::getValue)
.collect(Collectors.toSet());
var newFalseRecommendationsValues = newEntries.getFalseRecommendations()
.stream()
.map(DictionaryEntry::getValue)
.collect(Collectors.toSet());
oldModel.ifPresent(oldDictionaryModel -> {
oldModel.ifPresent(oldDictionaryModel -> {
});
// add old entries from existing DictionaryModel
oldModel.ifPresent(dictionaryModel -> entries.addAll(dictionaryModel.getEntries()
.stream()
.filter(f -> !newValues.contains(f.getValue()))
.collect(Collectors.toList())));
oldModel.ifPresent(dictionaryModel -> falsePositives.addAll(dictionaryModel.getFalsePositives()
.stream()
.filter(f -> !newFalsePositivesValues.contains(f.getValue()))
.collect(Collectors.toList())));
oldModel.ifPresent(dictionaryModel -> falseRecommendations.addAll(dictionaryModel.getFalseRecommendations()
.stream()
.filter(f -> !newFalseRecommendationsValues.contains(f.getValue()))
.collect(Collectors.toList())));
});
// add old entries from existing DictionaryModel
oldModel.ifPresent(dictionaryModel -> entries.addAll(dictionaryModel.getEntries().stream().filter(
f -> !newValues.contains(f.getValue())).collect(Collectors.toList())
));
oldModel.ifPresent(dictionaryModel -> falsePositives.addAll(dictionaryModel.getFalsePositives().stream().filter(
f -> !newFalsePositivesValues.contains(f.getValue())).collect(Collectors.toList())
));
oldModel.ifPresent(dictionaryModel -> falseRecommendations.addAll(dictionaryModel.getFalseRecommendations().stream().filter(
f -> !newFalseRecommendationsValues.contains(f.getValue())).collect(Collectors.toList())
));
// Add Increments
entries.addAll(newEntries.getEntries());
falsePositives.addAll(newEntries.getFalsePositives());
falseRecommendations.addAll(newEntries.getFalseRecommendations());
// Add Increments
entries.addAll(newEntries.getEntries());
falsePositives.addAll(newEntries.getFalsePositives());
falseRecommendations.addAll(newEntries.getFalseRecommendations());
return new DictionaryModel(t.getType(), t.getRank(), convertColor(t.getHexColor()), t.isCaseInsensitive(), t
.isHint(), entries, falsePositives, falseRecommendations, dossierId != null);
})
.sorted(Comparator.comparingInt(DictionaryModel::getRank).reversed())
.collect(Collectors.toList());
return new DictionaryModel(t.getType(), t.getRank(), convertColor(t.getHexColor()), t.isCaseInsensitive(), t.isHint(), entries, falsePositives, falseRecommendations, dossierId != null);
}).sorted(Comparator.comparingInt(DictionaryModel::getRank).reversed()).collect(Collectors.toList());
dictionary.forEach(dm -> dictionaryRepresentation.getLocalAccessMap().put(dm.getType(), dm));
@@ -193,11 +212,11 @@ public class DictionaryService {
falsePositives.forEach(entry -> entry.setValue(entry.getValue().toLowerCase(Locale.ROOT)));
falseRecommendations.forEach(entry -> entry.setValue(entry.getValue().toLowerCase(Locale.ROOT)));
}
log.info("Dictionary update returned {} entries {} falsePositives and {} falseRecommendations for type {}", entries.size(), falsePositives.size(), falseRecommendations.size(), type.getType());
return new DictionaryEntries(entries, falsePositives, falseRecommendations);
}
private float[] convertColor(String hex) {
Color color = Color.decode(hex);
@@ -225,6 +244,7 @@ public class DictionaryService {
}
@Timed("redactmanager_getDeepCopyDictionary")
public Dictionary getDeepCopyDictionary(String dossierTemplateId, String dossierId) {
List<DictionaryModel> copy = new ArrayList<>();
@@ -244,7 +264,12 @@ public class DictionaryService {
dossierDictionaryVersion = dossierRepresentation.getDictionaryVersion();
}
return new Dictionary(copy.stream().sorted(Comparator.comparingInt(DictionaryModel::getRank).reversed()).collect(Collectors.toList()), DictionaryVersion.builder().dossierTemplateVersion(dossierTemplateRepresentation.getDictionaryVersion()).dossierVersion(dossierDictionaryVersion).build());
return new Dictionary(copy.stream()
.sorted(Comparator.comparingInt(DictionaryModel::getRank).reversed())
.collect(Collectors.toList()), DictionaryVersion.builder()
.dossierTemplateVersion(dossierTemplateRepresentation.getDictionaryVersion())
.dossierVersion(dossierDictionaryVersion)
.build());
}
@@ -255,10 +280,12 @@ public class DictionaryService {
private Long getVersion(DictionaryRepresentation dictionaryRepresentation) {
if (dictionaryRepresentation == null) {
return null;
} else {
return dictionaryRepresentation.getDictionaryVersion();
}
}
}
@@ -3,6 +3,8 @@ package com.iqser.red.service.redaction.v1.server.redaction.service;
import com.iqser.red.service.redaction.v1.server.client.RulesClient;
import com.iqser.red.service.redaction.v1.server.exception.RulesValidationException;
import com.iqser.red.service.redaction.v1.server.redaction.model.Section;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import org.apache.commons.lang3.StringUtils;
import org.kie.api.KieServices;
@@ -41,6 +43,7 @@ public class DroolsExecutionService {
}
@Timed("redactmanager_executeRules")
public Section executeRules(KieContainer kieContainer, Section section) {
KieSession kieSession = kieContainer.newKieSession();
@@ -1,6 +1,20 @@
package com.iqser.red.service.redaction.v1.server.redaction.service;
import java.util.ArrayList;
import java.util.HashMap;
import java.util.HashSet;
import java.util.List;
import java.util.Map;
import java.util.Set;
import java.util.stream.Collectors;
import java.util.stream.Stream;
import org.apache.commons.lang3.StringUtils;
import org.kie.api.runtime.KieContainer;
import org.springframework.stereotype.Service;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.AnnotationStatus;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.ManualRedactions;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.IdRemoval;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualImageRecategorization;
import com.iqser.red.service.redaction.v1.model.AnalyzeRequest;
@@ -8,20 +22,24 @@ import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.server.classification.model.SectionText;
import com.iqser.red.service.redaction.v1.server.client.model.NerEntities;
import com.iqser.red.service.redaction.v1.server.redaction.model.Dictionary;
import com.iqser.red.service.redaction.v1.server.redaction.model.*;
import com.iqser.red.service.redaction.v1.server.redaction.model.DictionaryModel;
import com.iqser.red.service.redaction.v1.server.redaction.model.Entities;
import com.iqser.red.service.redaction.v1.server.redaction.model.Entity;
import com.iqser.red.service.redaction.v1.server.redaction.model.EntityPositionSequence;
import com.iqser.red.service.redaction.v1.server.redaction.model.EntityType;
import com.iqser.red.service.redaction.v1.server.redaction.model.Image;
import com.iqser.red.service.redaction.v1.server.redaction.model.PageEntities;
import com.iqser.red.service.redaction.v1.server.redaction.model.SearchableText;
import com.iqser.red.service.redaction.v1.server.redaction.model.Section;
import com.iqser.red.service.redaction.v1.server.redaction.model.SectionSearchableTextPair;
import com.iqser.red.service.redaction.v1.server.redaction.utils.EntitySearchUtils;
import com.iqser.red.service.redaction.v1.server.redaction.utils.FindEntityDetails;
import com.iqser.red.service.redaction.v1.server.redaction.utils.IdBuilder;
import com.iqser.red.service.redaction.v1.server.settings.RedactionServiceSettings;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.apache.commons.lang3.StringUtils;
import org.kie.api.runtime.KieContainer;
import org.springframework.stereotype.Service;
import java.util.*;
import java.util.stream.Collectors;
import java.util.stream.Stream;
@Slf4j
@Service
@@ -33,6 +51,7 @@ public class EntityRedactionService {
private final SurroundingWordsService surroundingWordsService;
public PageEntities findEntities(Dictionary dictionary, List<SectionText> sectionTexts, KieContainer kieContainer,
AnalyzeRequest analyzeRequest, NerEntities nerEntities) {
@@ -61,8 +80,7 @@ public class EntityRedactionService {
List<SectionSearchableTextPair> sectionSearchableTextPairs = new ArrayList<>();
for (SectionText reanalysisSection : reanalysisSections) {
Entities entities = findEntities(reanalysisSection.getSearchableText(), reanalysisSection.getHeadline(), reanalysisSection
.getSectionNumber(), dictionary, local, nerEntities, reanalysisSection.getCellStarts());
Entities entities = findEntities(reanalysisSection.getSearchableText(), reanalysisSection.getHeadline(), reanalysisSection.getSectionNumber(), dictionary, local, nerEntities, reanalysisSection.getCellStarts(), analyzeRequest.getManualRedactions());
if (reanalysisSection.getCellStarts() != null && !reanalysisSection.getCellStarts().isEmpty()) {
surroundingWordsService.addSurroundingText(entities.getEntities(), reanalysisSection.getSearchableText(), dictionary, reanalysisSection
@@ -71,19 +89,26 @@ public class EntityRedactionService {
surroundingWordsService.addSurroundingText(entities.getEntities(), reanalysisSection.getSearchableText(), dictionary);
}
if (!local && analyzeRequest.getManualRedactions() != null) {
var approvedForceRedactions = analyzeRequest.getManualRedactions().getForceRedactions().stream()
var approvedForceRedactions = analyzeRequest.getManualRedactions()
.getForceRedactions()
.stream()
.filter(fr -> fr.getStatus() == AnnotationStatus.APPROVED)
.filter(fr -> fr.getRequestDate() != null)
.collect(Collectors.toList());
// only approved id removals, that haven't been forced back afterwards
var idsToRemove = analyzeRequest.getManualRedactions().getIdsToRemove().stream()
var idsToRemove = analyzeRequest.getManualRedactions()
.getIdsToRemove()
.stream()
.filter(idr -> idr.getStatus() == AnnotationStatus.APPROVED && !idr.isRemoveFromDictionary())
.filter(idr -> idr.getRequestDate() != null)
.filter(idr -> approvedForceRedactions.stream().noneMatch(forceRedact -> forceRedact.getRequestDate().isAfter(idr.getRequestDate())))
.map(IdRemoval::getAnnotationId).collect(Collectors.toSet());
.filter(idr -> approvedForceRedactions.stream()
.noneMatch(forceRedact -> forceRedact.getAnnotationId()
.equals(idr.getAnnotationId()) && forceRedact.getRequestDate()
.isAfter(idr.getRequestDate())))
.map(IdRemoval::getAnnotationId)
.collect(Collectors.toSet());
if (reanalysisSection.getImages() != null && !reanalysisSection.getImages()
.isEmpty() && analyzeRequest.getManualRedactions().getImageRecategorization() != null) {
@@ -110,13 +135,13 @@ public class EntityRedactionService {
}));
}
log.debug("Section {}, Images: {}", reanalysisSection.getSectionNumber(), reanalysisSection.getImages());
sectionSearchableTextPairs.add(new SectionSearchableTextPair(Section.builder()
.isLocal(false)
.dictionaryTypes(dictionary.getTypes())
.entities(hintsPerSectionNumber != null && hintsPerSectionNumber.containsKey(reanalysisSection.getSectionNumber()) ? Stream
.concat(entities.getEntities().stream(), hintsPerSectionNumber.get(reanalysisSection.getSectionNumber())
.stream())
.entities(hintsPerSectionNumber != null && hintsPerSectionNumber.containsKey(reanalysisSection.getSectionNumber()) ? Stream.concat(entities.getEntities()
.stream(), hintsPerSectionNumber.get(reanalysisSection.getSectionNumber()).stream())
.collect(Collectors.toSet()) : entities.getEntities())
.nerEntities(entities.getNerEntities())
.text(reanalysisSection.getSearchableText().getAsStringWithLinebreaks())
@@ -127,8 +152,10 @@ public class EntityRedactionService {
.searchableText(reanalysisSection.getSearchableText())
.dictionary(dictionary)
.images(reanalysisSection.getImages())
.sectionAreas(reanalysisSection.getSectionAreas())
.fileAttributes(analyzeRequest.getFileAttributes())
.build(), reanalysisSection.getSearchableText()));
.manualRedactions(analyzeRequest.getManualRedactions())
.build(), reanalysisSection.getSearchableText(), reanalysisSection.getCellStarts()));
}
@@ -136,6 +163,16 @@ public class EntityRedactionService {
sectionSearchableTextPairs.forEach(sectionSearchableTextPair -> {
Section analysedSection = droolsExecutionService.executeRules(kieContainer, sectionSearchableTextPair.getSection());
EntitySearchUtils.removeEntitiesContainedInLarger(analysedSection.getEntities());
var entriesWithoutSurroundingText = analysedSection.getEntities().stream().filter(e -> e.getTextAfter() == null && e.getTextBefore() == null).collect(Collectors.toSet());
if (sectionSearchableTextPair.getCellStarts() != null && !sectionSearchableTextPair.getCellStarts()
.isEmpty()) {
surroundingWordsService.addSurroundingText(entriesWithoutSurroundingText, sectionSearchableTextPair.getSearchableText(), dictionary, sectionSearchableTextPair.getCellStarts());
} else {
surroundingWordsService.addSurroundingText(entriesWithoutSurroundingText, sectionSearchableTextPair.getSearchableText(), dictionary);
}
entities.addAll(analysedSection.getEntities());
if (!local) {
@@ -144,6 +181,7 @@ public class EntityRedactionService {
}
addLocalValuesToDictionary(analysedSection, dictionary);
}
});
return entities;
@@ -162,11 +200,7 @@ public class EntityRedactionService {
for (Map.Entry<Integer, List<EntityPositionSequence>> entry : sequenceOnPage.entrySet()) {
entitiesPerPage.computeIfAbsent(entry.getKey(), (x) -> new ArrayList<>())
.add(new Entity(entity.getWord(), entity.getType(), entity.isRedaction(), entity.getRedactionReason(), entry
.getValue(), entity.getHeadline(), entity.getMatchedRule(), entity.getSectionNumber(), entity
.getLegalBasis(), entity.isDictionaryEntry(), entity.getTextBefore(), entity.getTextAfter(), entity
.getStart(), entity.getEnd(), entity.isDossierDictionaryEntry(), entity.getEngines(), entity
.getReferences(), entity.getEntityType()));
.add(new Entity(entity.getWord(), entity.getType(), entity.isRedaction(), entity.getRedactionReason(), entry.getValue(), entity.getHeadline(), entity.getMatchedRule(), entity.getSectionNumber(), entity.getLegalBasis(), entity.isDictionaryEntry(), entity.getTextBefore(), entity.getTextAfter(), entity.getStart(), entity.getEnd(), entity.isDossierDictionaryEntry(), entity.getEngines(), entity.getReferences(), entity.getEntityType()));
}
}
return entitiesPerPage;
@@ -204,9 +238,10 @@ public class EntityRedactionService {
}
@Timed("redactmanager_findEntities")
private Entities findEntities(SearchableText searchableText, String headline, int sectionNumber,
Dictionary dictionary, boolean local, NerEntities nerEntities,
List<Integer> cellStarts) {
List<Integer> cellStarts, ManualRedactions manualRedactions) {
Set<Entity> found = new HashSet<>();
String searchableString = searchableText.asString();
@@ -219,8 +254,7 @@ public class EntityRedactionService {
for (DictionaryModel model : dictionary.getDictionaryModels()) {
var searchImplementation = local ? model.getLocalSearch() : model.getEntriesSearch();
var entities = EntitySearchUtils.findEntities(model.isCaseInsensitive() ? lowercaseInputString : searchableString,
searchImplementation, model, new FindEntityDetails(model.getType(),headline, sectionNumber, !local, model.isDossierDictionary(), local ? Engine.RULE : Engine.DICTIONARY, local? EntityType.RECOMMENDATION: EntityType.ENTITY));
var entities = EntitySearchUtils.findEntities(model.isCaseInsensitive() ? lowercaseInputString : searchableString, searchImplementation, model, new FindEntityDetails(model.getType(), headline, sectionNumber, !local, model.isDossierDictionary(), local ? Engine.RULE : Engine.DICTIONARY, local ? EntityType.RECOMMENDATION : EntityType.ENTITY));
EntitySearchUtils.addOrAddEngine(found, entities);
}
@@ -230,7 +264,7 @@ public class EntityRedactionService {
nerFound.addAll(getNerValues(sectionNumber, nerEntities, cellStarts, headline));
}
return new Entities(EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary), nerFound);
return new Entities(EntitySearchUtils.clearAndFindPositions(found, searchableText, dictionary, manualRedactions), nerFound);
}
@@ -7,10 +7,12 @@ import org.springframework.stereotype.Service;
import com.iqser.red.service.redaction.v1.model.ImportedRedaction;
import com.iqser.red.service.redaction.v1.model.ImportedRedactions;
import com.iqser.red.service.redaction.v1.model.Point;
import com.iqser.red.service.redaction.v1.model.Rectangle;
import com.iqser.red.service.redaction.v1.model.RedactionLogEntry;
import com.iqser.red.service.redaction.v1.server.storage.RedactionStorageService;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
@Service
@@ -23,6 +25,7 @@ public class ImportedRedactionService {
private final RedactionStorageService redactionStorageService;
@Timed("redactmanager_processImportedRedactions")
public List<RedactionLogEntry> processImportedRedactions(String dossierTemplateId, String dossierId, String fileId,
List<RedactionLogEntry> redactionLogEntries,
boolean addImportedRedactions) {
@@ -68,12 +71,13 @@ public class ImportedRedactionService {
private void addIntersections(RedactionLogEntry redactionLogEntry, ImportedRedactions importedRedactions) {
for (Rectangle rectangle : redactionLogEntry.getPositions()) {
var normalizedRectangle = normalize(rectangle);
if (importedRedactions.getImportedRedactions().containsKey(rectangle.getPage())) {
var importedRedactionsOnPage = importedRedactions.getImportedRedactions().get(rectangle.getPage());
for (ImportedRedaction importedRedaction : importedRedactionsOnPage) {
for (Rectangle importedRedactionPosition : importedRedaction.getPositions()) {
if (intersects(importedRedactionPosition, rectangle)) {
if(redactionLogEntry.getImportedRedactionIntersections() == null){
if (rectOverlap(normalizedRectangle, normalize(importedRedactionPosition))) {
if (redactionLogEntry.getImportedRedactionIntersections() == null) {
redactionLogEntry.setImportedRedactionIntersections(new HashSet<>());
}
redactionLogEntry.getImportedRedactionIntersections().add(importedRedaction.getId());
@@ -85,15 +89,56 @@ public class ImportedRedactionService {
}
private boolean intersects(Rectangle importedPosition, Rectangle redactionLogPosition) {
boolean valueInRange(float value, float min, float max) {
return redactionLogPosition.getTopLeft()
.getX() + redactionLogPosition.getWidth() > importedPosition.getTopLeft()
.getX() && redactionLogPosition.getTopLeft()
.getY() + redactionLogPosition.getHeight() > importedPosition.getTopLeft()
.getY() && redactionLogPosition.getTopLeft().getX() < importedPosition.getTopLeft()
.getX() + importedPosition.getWidth() && redactionLogPosition.getTopLeft()
.getY() < importedPosition.getTopLeft().getY() + importedPosition.getHeight();
return round(value) >= round(min) && round(value) <= round(max);
}
boolean rectOverlap(Rectangle a, Rectangle b) {
boolean xOverlap = valueInRange(a.getTopLeft().getX(), b.getTopLeft().getX(), b.getTopLeft()
.getX() + b.getWidth()) || valueInRange(b.getTopLeft().getX(), a.getTopLeft().getX(), a.getTopLeft()
.getX() + a.getWidth());
boolean yOverlap = valueInRange(a.getTopLeft().getY(), b.getTopLeft().getY(), b.getTopLeft()
.getY() + b.getHeight()) || valueInRange(b.getTopLeft().getY(), a.getTopLeft().getY(), a.getTopLeft()
.getY() + a.getHeight());
return xOverlap && yOverlap;
}
private Rectangle normalize(Rectangle rectangle) {
Rectangle r = new Rectangle();
Point p = new Point();
if (rectangle.getWidth() < 0) {
p.setX(rectangle.getTopLeft().getX() - rectangle.getWidth());
r.setWidth(Math.abs(rectangle.getWidth()));
} else {
p.setX(rectangle.getTopLeft().getX());
r.setWidth(rectangle.getWidth());
}
if (rectangle.getHeight() < 0) {
p.setY(rectangle.getTopLeft().getY() + rectangle.getHeight());
r.setHeight(Math.abs(rectangle.getHeight()));
} else {
p.setY(rectangle.getTopLeft().getY());
}
r.setTopLeft(p);
r.setPage(rectangle.getPage());
return r;
}
private float round(float value) {
double d = Math.pow(10, 0);
return (float) (Math.round(value * d) / d);
}
@@ -1,9 +1,15 @@
package com.iqser.red.service.redaction.v1.server.redaction.service;
import java.util.ArrayList;
import java.util.List;
import java.util.Set;
import org.apache.commons.lang3.tuple.Pair;
import org.springframework.stereotype.Service;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.ManualRedactions;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.Rectangle;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualRedactionEntry;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.entitymapped.ManualResizeRedaction;
import com.iqser.red.service.redaction.v1.model.AnalyzeResult;
import com.iqser.red.service.redaction.v1.model.Engine;
import com.iqser.red.service.redaction.v1.model.SectionArea;
@@ -17,14 +23,10 @@ import com.iqser.red.service.redaction.v1.server.redaction.utils.EntitySearchUti
import com.iqser.red.service.redaction.v1.server.redaction.utils.FindEntityDetails;
import com.iqser.red.service.redaction.v1.server.redaction.utils.SearchImplementation;
import com.iqser.red.service.redaction.v1.server.storage.RedactionStorageService;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.apache.commons.lang3.tuple.Pair;
import org.springframework.stereotype.Service;
import java.util.ArrayList;
import java.util.List;
import java.util.Set;
@Slf4j
@Service
@@ -34,17 +36,16 @@ public class ManualRedactionSurroundingTextService {
private final RedactionStorageService redactionStorageService;
private final SurroundingWordsService surroundingWordsService;
@Timed("redactmanager_SurroundingTextAnalysis")
public AnalyzeResult addSurroundingText(String dossierId, String fileId, ManualRedactions manualRedactions) {
long startTime = System.currentTimeMillis();
Text text = redactionStorageService.getText(dossierId, fileId);
List<ManualRedactionEntry> processedAddRedactions = new ArrayList<>();
List<ManualResizeRedaction> processedResizeRedactions = new ArrayList<>();
for (SectionText sectionText : text.getSectionTexts()) {
if (manualRedactions.getEntriesToAdd().isEmpty() && manualRedactions.getResizeRedactions().isEmpty()) {
if (manualRedactions.getEntriesToAdd().isEmpty()) {
break;
}
@@ -61,23 +62,10 @@ public class ManualRedactionSurroundingTextService {
addItty.remove();
}
}
var resizeItty = manualRedactions.getResizeRedactions().iterator();
while (resizeItty.hasNext()) {
var manualResizeRedaction = resizeItty.next();
if (sectionContainsEntry(sectionArea, manualResizeRedaction.getPositions())) {
var surroundingText = findSurroundingText(sectionText, manualResizeRedaction.getValue(), manualResizeRedaction.getPositions());
manualResizeRedaction.setTextBefore(surroundingText.getLeft());
manualResizeRedaction.setTextAfter(surroundingText.getRight());
processedResizeRedactions.add(manualResizeRedaction);
resizeItty.remove();
}
}
}
}
manualRedactions.getEntriesToAdd().addAll(processedAddRedactions);
manualRedactions.getResizeRedactions().addAll(processedResizeRedactions);
return AnalyzeResult.builder()
.dossierId(dossierId)
@@ -88,11 +76,11 @@ public class ManualRedactionSurroundingTextService {
}
private Pair<String, String> findSurroundingText(SectionText sectionText, String value, List<Rectangle> toFindPositions) {
private Pair<String, String> findSurroundingText(SectionText sectionText, String value,
List<Rectangle> toFindPositions) {
Set<Entity> entities = EntitySearchUtils.find(sectionText.getText(), new SearchImplementation(value, false),
new FindEntityDetails("dummy", sectionText.getHeadline(), sectionText.getSectionNumber(), false, false, Engine.DICTIONARY, EntityType.ENTITY));
Set<Entity> entitiesWithPositions = EntitySearchUtils.clearAndFindPositions(entities, sectionText.getSearchableText(), null);
Set<Entity> entities = EntitySearchUtils.find(sectionText.getText(), new SearchImplementation(value, false), new FindEntityDetails("dummy", sectionText.getHeadline(), sectionText.getSectionNumber(), false, false, Engine.DICTIONARY, EntityType.ENTITY));
Set<Entity> entitiesWithPositions = EntitySearchUtils.clearAndFindPositions(entities, sectionText.getSearchableText(), null, null);
Entity correctEntity = getEntityOnCorrectPosition(entitiesWithPositions, toFindPositions);
@@ -2,6 +2,8 @@ package com.iqser.red.service.redaction.v1.server.redaction.service;
import com.iqser.red.service.redaction.v1.model.*;
import com.iqser.red.service.redaction.v1.server.storage.RedactionStorageService;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.springframework.stereotype.Service;
@@ -18,6 +20,7 @@ public class RedactionChangeLogService {
private final RedactionStorageService redactionStorageService;
@Timed("redactmanager_computeChanges")
public RedactionLogChanges computeChanges(String dossierId, String fileId, RedactionLog currentRedactionLog, int analysisNumber) {
long start = System.currentTimeMillis();
@@ -22,6 +22,7 @@ import com.iqser.red.service.redaction.v1.server.redaction.model.Image;
import com.iqser.red.service.redaction.v1.server.redaction.model.PageEntities;
import com.iqser.red.service.redaction.v1.server.redaction.utils.IdBuilder;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
@@ -33,6 +34,7 @@ public class RedactionLogCreatorService {
private final DictionaryService dictionaryService;
@Timed("redactmanager_createRedactionLog")
public List<RedactionLogEntry> createRedactionLog(PageEntities pageEntities, int numberOfPages,
String dossierTemplateId) {
@@ -12,6 +12,8 @@ import com.iqser.red.service.redaction.v1.model.Rectangle;
import com.iqser.red.service.redaction.v1.model.*;
import com.iqser.red.service.redaction.v1.server.exception.NotFoundException;
import com.iqser.red.service.redaction.v1.server.storage.RedactionStorageService;
import io.micrometer.core.annotation.Timed;
import lombok.AllArgsConstructor;
import lombok.Data;
import lombok.RequiredArgsConstructor;
@@ -34,6 +36,8 @@ public class RedactionLogMergeService {
private final RedactionStorageService redactionStorageService;
@Timed("redactmanager_getMergedRedactionLog")
public RedactionLog provideRedactionLog(RedactionRequest redactionRequest) {
log.debug("Requested preview for: {}", redactionRequest);
@@ -290,15 +294,23 @@ public class RedactionLogMergeService {
redactionLogEntry.setColor(getColor(redactionLogEntry.getType(), colors, false, redactionLogEntry.isRedacted(), false, types));
redactionLogEntry.setPositions(convertPositions(manualResizeRedact.getPositions()));
redactionLogEntry.setValue(manualResizeRedact.getValue());
redactionLogEntry.setTextBefore(manualResizeRedact.getTextBefore());
redactionLogEntry.setTextAfter(manualResizeRedact.getTextAfter());
// This is for backwards compatibility, now the text after/before is calculated during reanalysis because we need to find dict entries on positions where entries are resized to smaller.
if(manualResizeRedact.getTextBefore() != null || manualResizeRedact.getTextAfter() != null) {
redactionLogEntry.setTextBefore(manualResizeRedact.getTextBefore());
redactionLogEntry.setTextAfter(manualResizeRedact.getTextAfter());
}
manualOverrideReason = mergeReasonIfNecessary(redactionLogEntry.getReason(), ", resized by manual override");
} else if (manualResizeRedact.getStatus().equals(AnnotationStatus.REQUESTED)) {
manualOverrideReason = mergeReasonIfNecessary(redactionLogEntry.getReason(), ", requested to resize redact");
redactionLogEntry.setColor(getColor(redactionLogEntry.getType(), colors, true, redactionLogEntry.isRedacted(), false, types));
redactionLogEntry.setPositions(convertPositions(manualResizeRedact.getPositions()));
redactionLogEntry.setTextBefore(manualResizeRedact.getTextBefore());
redactionLogEntry.setTextAfter(manualResizeRedact.getTextAfter());
// This is for backwards compatibility, now the text after/before is calculated during reanalysis because we need to find dict entries on positions where entries are resized to smaller.
if(manualResizeRedact.getTextBefore() != null || manualResizeRedact.getTextAfter() != null) {
redactionLogEntry.setTextBefore(manualResizeRedact.getTextBefore());
redactionLogEntry.setTextAfter(manualResizeRedact.getTextAfter());
}
}
redactionLogEntry.setReason(manualOverrideReason);
@@ -4,6 +4,8 @@ import com.iqser.red.service.redaction.v1.server.redaction.model.Dictionary;
import com.iqser.red.service.redaction.v1.server.redaction.model.Entity;
import com.iqser.red.service.redaction.v1.server.redaction.model.SearchableText;
import com.iqser.red.service.redaction.v1.server.settings.RedactionServiceSettings;
import io.micrometer.core.annotation.Timed;
import lombok.RequiredArgsConstructor;
import lombok.extern.slf4j.Slf4j;
import org.springframework.stereotype.Service;
@@ -19,6 +21,7 @@ public class SurroundingWordsService {
private final RedactionServiceSettings redactionServiceSettings;
@Timed("redactmanager_addSurroundingText")
public void addSurroundingText(Set<Entity> entities, SearchableText searchableText, Dictionary dictionary) {
if (entities.isEmpty()) {
@@ -39,6 +42,7 @@ public class SurroundingWordsService {
}
@Timed("redactmanager_addSurroundingText_tables")
public void addSurroundingText(Set<Entity> entities, SearchableText searchableText, Dictionary dictionary,
List<Integer> cellstarts) {
@@ -1,7 +1,11 @@
package com.iqser.red.service.redaction.v1.server.redaction.utils;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.AnnotationStatus;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.ManualRedactions;
import com.iqser.red.service.redaction.v1.server.redaction.model.Dictionary;
import com.iqser.red.service.redaction.v1.server.redaction.model.*;
import io.micrometer.core.annotation.Timed;
import lombok.experimental.UtilityClass;
import lombok.extern.slf4j.Slf4j;
@@ -55,7 +59,7 @@ public class EntitySearchUtils {
}
public Set<Entity> clearAndFindPositions(Set<Entity> entities, SearchableText text, Dictionary dictionary) {
public Set<Entity> clearAndFindPositions(Set<Entity> entities, SearchableText text, Dictionary dictionary, ManualRedactions manualRedactions) {
Map<String, List<Entity>> entitiesByWord = new HashMap<>();
@@ -79,12 +83,46 @@ public class EntitySearchUtils {
}
}
if (manualRedactions != null && manualRedactions.getResizeRedactions() != null && !manualRedactions.getResizeRedactions().isEmpty()){
applyResizeRedactions(entities, manualRedactions);
}
removeEntitiesContainedInLarger(entities);
return entities;
}
private Set<Entity> applyResizeRedactions(Set<Entity> entitiesWithPositions, ManualRedactions manualRedactions) {
if (manualRedactions == null || manualRedactions.getResizeRedactions() == null || manualRedactions.getResizeRedactions().isEmpty()){
return entitiesWithPositions;
}
entitiesWithPositions.forEach(e -> e.getPositionSequences().forEach(pos -> {
manualRedactions.getResizeRedactions().stream().filter(resize -> resize.getStatus().equals(AnnotationStatus.APPROVED)).forEach(resize -> {
if (resize.getAnnotationId().equals(pos.getId())) {
if (resize.getValue().length() < e.getWord().length() && e.getWord().contains(resize.getValue())) {
int start = e.getWord().indexOf(resize.getValue());
e.setStart(e.getStart() + start);
e.setEnd(e.getStart() + resize.getValue().length());
e.setResized(true);
e.setWord(resize.getValue());
} else if(resize.getValue().length() > e.getWord().length() && resize.getValue().contains(e.getWord())){
int start = resize.getValue().indexOf(e.getWord());
e.setStart(e.getStart() - start);
e.setEnd(e.getStart() + resize.getValue().length());
e.setResized(true);
e.setWord(resize.getValue());
}
}
});
}));
return entitiesWithPositions;
}
public void removeFalsePositives(Set<Entity> entities, Set<Entity> falsePositives) {
List<Entity> wordsToRemove = new ArrayList<>();
@@ -112,7 +150,11 @@ public class EntitySearchUtils {
wordsToRemove.add(word);
} else if(!(inner.getEntityType() == EntityType.FALSE_RECOMMENDATION && word.getEntityType() == EntityType.ENTITY ||
inner.getEntityType() == EntityType.ENTITY && word.getEntityType() == EntityType.FALSE_RECOMMENDATION)) {
wordsToRemove.add(inner);
if(inner.isResized()){
wordsToRemove.add(word);
} else {
wordsToRemove.add(inner);
}
}
}
}
@@ -166,6 +208,9 @@ public class EntitySearchUtils {
if (existing.getType().equals(found.getType())) {
existing.getEngines().addAll(found.getEngines());
existing.setLegalBasis(found.getLegalBasis());
existing.setMatchedRule(found.getMatchedRule());
existing.setRedactionReason(found.getRedactionReason());
if (existing.getEntityType().equals(EntityType.RECOMMENDATION) && found.getEntityType().equals(EntityType.ENTITY)
|| existing.getEntityType().equals(EntityType.ENTITY) && found.getEntityType().equals(EntityType.RECOMMENDATION)) {
existing.setEntityType(EntityType.ENTITY);
@@ -230,6 +275,9 @@ public class EntitySearchUtils {
}
var existingEntity = existingOptional.get();
existingEntity.getEngines().addAll(toAdd.getEngines());
existingEntity.setLegalBasis(toAdd.getLegalBasis());
existingEntity.setMatchedRule(toAdd.getMatchedRule());
existingEntity.setRedactionReason(toAdd.getRedactionReason());
} else {
existing.add(toAdd);
}
@@ -39,7 +39,7 @@ public class ImageService {
imageServiceResponse.getData().forEach(imageMetadata -> {
var classification = imageMetadata.getFilters().isAllPassed() ? ImageType.valueOf(imageMetadata.getClassification().getLabel().toUpperCase(Locale.ROOT)) : ImageType.OTHER;
images.computeIfAbsent(imageMetadata.getPosition().getPageNumber() ,x -> new ArrayList<>())
.add(new PdfImage(new RedRectangle2D(imageMetadata.getPosition().getX1(), imageMetadata.getPosition().getY1(), imageMetadata.getGeometry().getWidth(), imageMetadata.getGeometry().getHeight()), classification, imageMetadata.getPosition().getPageNumber()));
.add(new PdfImage(new RedRectangle2D(imageMetadata.getPosition().getX1(), imageMetadata.getPosition().getY1(), imageMetadata.getGeometry().getWidth(), imageMetadata.getGeometry().getHeight()), classification,imageMetadata.isAlpha(), imageMetadata.getPosition().getPageNumber()));
});
return images;
@@ -1,16 +1,31 @@
package com.iqser.red.service.redaction.v1.server.segmentation;
import com.iqser.red.service.redaction.v1.server.classification.model.*;
import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.Iterator;
import java.util.List;
import java.util.Map;
import java.util.stream.Collectors;
import org.apache.commons.collections4.CollectionUtils;
import org.springframework.stereotype.Service;
import com.iqser.red.service.redaction.v1.server.classification.model.Document;
import com.iqser.red.service.redaction.v1.server.classification.model.Footer;
import com.iqser.red.service.redaction.v1.server.classification.model.Header;
import com.iqser.red.service.redaction.v1.server.classification.model.Page;
import com.iqser.red.service.redaction.v1.server.classification.model.Paragraph;
import com.iqser.red.service.redaction.v1.server.classification.model.TextBlock;
import com.iqser.red.service.redaction.v1.server.classification.model.UnclassifiedText;
import com.iqser.red.service.redaction.v1.server.redaction.model.PdfImage;
import com.iqser.red.service.redaction.v1.server.tableextraction.model.AbstractTextContainer;
import com.iqser.red.service.redaction.v1.server.tableextraction.model.Cell;
import com.iqser.red.service.redaction.v1.server.tableextraction.model.Table;
import org.apache.commons.collections4.CollectionUtils;
import org.springframework.stereotype.Service;
import java.util.*;
import java.util.stream.Collectors;
import lombok.extern.slf4j.Slf4j;
@Slf4j
@Service
public class SectionsBuilderService {
@@ -53,8 +68,7 @@ public class SectionsBuilderService {
continue;
}
if (prev != null && current.getClassification().startsWith("H ") && !prev.getClassification()
.startsWith("H ") || !document.isHeadlines()) {
if (prev != null && current.getClassification().startsWith("H ") && !prev.getClassification().startsWith("H ") || !document.isHeadlines()) {
Paragraph chunkBlock = buildTextBlock(chunkWords, lastHeadline);
chunkBlock.setHeadline(lastHeadline);
if (document.isHeadlines()) {
@@ -100,17 +114,19 @@ public class SectionsBuilderService {
public void addImagesToSections(Document document) {
Map<Integer, SortedSet<Paragraph>> paragraphMap = new HashMap<>();
Map<Integer, List<Paragraph>> paragraphMap = new HashMap<>();
for (Paragraph paragraph : document.getParagraphs()) {
for (AbstractTextContainer container : paragraph.getPageBlocks()) {
paragraphMap.computeIfAbsent(container.getPage(), x -> new TreeSet<>()).add(paragraph);
paragraphMap.computeIfAbsent(container.getPage(), c -> new ArrayList<>()).add(paragraph);
}
}
if (paragraphMap.isEmpty()) {
Paragraph paragraph = new Paragraph();
document.getParagraphs().add(paragraph);
paragraphMap.computeIfAbsent(1, x -> new TreeSet<>()).add(paragraph);
paragraphMap.computeIfAbsent(1, x -> new ArrayList<>()).add(paragraph);
}
// first page is always a paragraph, else we can't process pages 1..N,
@@ -118,12 +134,12 @@ public class SectionsBuilderService {
if (paragraphMap.get(1) == null) {
Paragraph paragraph = new Paragraph();
document.getParagraphs().add(paragraph);
paragraphMap.computeIfAbsent(1, x -> new TreeSet<>()).add(paragraph);
paragraphMap.computeIfAbsent(1, x -> new ArrayList<>()).add(paragraph);
}
for (Page page : document.getPages()) {
for (PdfImage image : page.getImages()) {
SortedSet<Paragraph> paragraphsOnPage = paragraphMap.get(page.getPageNumber());
List<Paragraph> paragraphsOnPage = paragraphMap.get(page.getPageNumber());
if (paragraphsOnPage == null) {
int i = page.getPageNumber();
while (paragraphsOnPage == null) {
@@ -131,27 +147,63 @@ public class SectionsBuilderService {
i--;
}
}
Float perviousEnd = 0f;
for (Paragraph paragraph : paragraphsOnPage) {
Float currentEnd = 0f;
Float xMin = null;
Float yMin = null;
Float xMax = null;
Float yMax = null;
for (AbstractTextContainer abs : paragraph.getPageBlocks()) {
if (abs.getPage() != page.getPageNumber()) {
continue;
}
if (abs.getMaxY() > currentEnd) {
currentEnd = abs.getMaxY();
if (abs.getMinX() < abs.getMaxX()) {
if (xMin == null || abs.getMinX() < xMin) {
xMin = abs.getMinX();
}
if (xMax == null || abs.getMaxX() > xMax) {
xMax = abs.getMaxX();
}
} else {
if (xMin == null || abs.getMaxX() < xMin) {
xMin = abs.getMaxX();
}
if (xMax == null || abs.getMinX() > xMax) {
xMax = abs.getMinX();
}
}
if (abs.getMinY() < abs.getMaxY()) {
if (yMin == null || abs.getMinY() < yMin) {
yMin = abs.getMinY();
}
if (yMax == null || abs.getMaxY() > yMax) {
yMax = abs.getMaxY();
}
} else {
if (yMin == null || abs.getMaxY() < yMin) {
yMin = abs.getMaxY();
}
if (yMax == null || abs.getMinY() > yMax) {
yMax = abs.getMinY();
}
}
}
if (image.getPosition().getY() >= perviousEnd && image.getPosition().getY() <= currentEnd) {
log.debug("Image position x: {}, y: {}", image.getPosition().getX(), image.getPosition().getY());
log.debug("Paragraph position xMin: {}, xMax: {}, yMin: {}, yMax: {}", xMin, xMax, yMin, yMax);
if (xMin != null && xMax != null && yMin != null && yMax != null && image.getPosition().getX() >= xMin && image.getPosition()
.getX() <= xMax && image.getPosition().getY() >= yMin && image.getPosition().getY() <= yMax) {
paragraph.getImages().add(image);
image.setAppendedToParagraph(true);
}
perviousEnd = currentEnd;
}
if (!image.isAppendedToParagraph()) {
paragraphsOnPage.first().getImages().add(image);
log.debug("Image uses first paragraph");
paragraphsOnPage.get(0).getImages().add(image);
image.setAppendedToParagraph(true);
}
}
@@ -166,9 +218,7 @@ public class SectionsBuilderService {
List<Cell> previousTableNonHeaderRow = getRowWithNonHeaderCells(previousTable);
List<Cell> tableNonHeaderRow = getRowWithNonHeaderCells(currentTable);
// Allow merging of tables if header row is separated from first logical non-header row
if (previousTableNonHeaderRow.isEmpty() && previousTable.getRowCount() == 1 && previousTable.getRows()
.get(0)
.size() == tableNonHeaderRow.size()) {
if (previousTableNonHeaderRow.isEmpty() && previousTable.getRowCount() == 1 && previousTable.getRows().get(0).size() == tableNonHeaderRow.size()) {
previousTableNonHeaderRow = previousTable.getRows().get(0).stream().map(cell -> {
Cell fakeCell = new Cell(cell.getPoints()[0], cell.getPoints()[2]);
fakeCell.setHeaderCells(Collections.singletonList(cell));
@@ -178,8 +228,7 @@ public class SectionsBuilderService {
if (previousTableNonHeaderRow.size() == tableNonHeaderRow.size()) {
for (int i = currentTable.getRowCount() - 1; i >= 0; i--) { // Non header rows are most likely at bottom of table
List<Cell> row = currentTable.getRows().get(i);
if (row.size() == tableNonHeaderRow.size() && row.stream()
.allMatch(cell -> cell.getHeaderCells().isEmpty())) {
if (row.size() == tableNonHeaderRow.size() && row.stream().allMatch(cell -> cell.getHeaderCells().isEmpty())) {
for (int j = 0; j < row.size(); j++) {
row.get(j).setHeaderCells(previousTableNonHeaderRow.get(j).getHeaderCells());
}
@@ -229,24 +278,20 @@ public class SectionsBuilderService {
TextBlock wordBlock = (TextBlock) container;
if (textBlock == null) {
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock
.getSequences(), wordBlock.getRotation());
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock.getSequences(), wordBlock.getRotation());
textBlock.setPage(wordBlock.getPage());
} else if (splitByTable) {
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock
.getSequences(), wordBlock.getRotation());
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock.getSequences(), wordBlock.getRotation());
textBlock.setPage(wordBlock.getPage());
alreadyAdded = false;
} else if (pageBefore != -1 && wordBlock.getPage() != pageBefore) {
textBlock.setPage(pageBefore);
paragraph.getPageBlocks().add(textBlock);
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock
.getSequences(), wordBlock.getRotation());
textBlock = new TextBlock(wordBlock.getMinX(), wordBlock.getMaxX(), wordBlock.getMinY(), wordBlock.getMaxY(), wordBlock.getSequences(), wordBlock.getRotation());
textBlock.setPage(wordBlock.getPage());
} else {
TextBlock spatialEntity = textBlock.union(wordBlock);
textBlock.resize(spatialEntity.getMinX(), spatialEntity.getMinY(), spatialEntity.getWidth(), spatialEntity
.getHeight());
textBlock.resize(spatialEntity.getMinX(), spatialEntity.getMinY(), spatialEntity.getWidth(), spatialEntity.getHeight());
}
pageBefore = wordBlock.getPage();
splitByTable = false;
@@ -268,11 +313,7 @@ public class SectionsBuilderService {
private boolean hasInvalidHeaderInformation(Table table) {
return table.getRows()
.stream()
.flatMap(row -> row.stream().filter(cell -> CollectionUtils.isNotEmpty(cell.getHeaderCells())))
.findAny()
.isEmpty();
return table.getRows().stream().flatMap(row -> row.stream().filter(cell -> CollectionUtils.isNotEmpty(cell.getHeaderCells()))).findAny().isEmpty();
}
@@ -12,10 +12,13 @@ import com.iqser.red.service.redaction.v1.server.client.model.NerEntities;
import com.iqser.red.service.redaction.v1.server.exception.NotFoundException;
import com.iqser.red.storage.commons.exception.StorageObjectDoesNotExist;
import com.iqser.red.storage.commons.service.StorageService;
import io.micrometer.core.annotation.Timed;
import lombok.Getter;
import lombok.RequiredArgsConstructor;
import lombok.SneakyThrows;
import lombok.extern.slf4j.Slf4j;
import org.springframework.core.io.InputStreamResource;
import org.springframework.stereotype.Service;
@@ -41,21 +44,6 @@ public class RedactionStorageService {
}
@SneakyThrows
public void storeObject(String dossierId, String fileId, FileType fileType, Object any) {
var baos = new ByteArrayOutputStream();
try {
dslJson.serialize(any, baos);
} catch (com.dslplatform.json.SerializationException e){
// Fails on file 49 Cyprodinil - EU AIR3 - MCA Section 8 Supplement - Ecotoxicological studies on the active substance.pdf
var bytes = objectMapper.writeValueAsBytes(any);
storageService.storeObject(StorageIdUtils.getStorageId(dossierId, fileId, fileType), bytes);
dslJson.newWriter();
return;
}
storageService.storeObject(StorageIdUtils.getStorageId(dossierId, fileId, fileType), baos.toByteArray());
}
@SneakyThrows
public void storeObject(String dossierId, String fileId, FileType fileType, InputStream inputStream) {
@@ -63,6 +51,27 @@ public class RedactionStorageService {
}
@SneakyThrows
@Timed("redactmanager_storeObject")
public void storeObject(String dossierId, String fileId, FileType fileType, Object any) {
var bytes = serializeObject(any);
storageService.storeObject(StorageIdUtils.getStorageId(dossierId, fileId, fileType), bytes);
}
@SneakyThrows
@Timed("redactmanager_serializeObject")
private byte[] serializeObject(Object any) {
var baos = new ByteArrayOutputStream();
dslJson.serialize(any, baos);
return baos.toByteArray();
}
@Timed("redactmanager_getImportedRedactions")
public ImportedRedactions getImportedRedactions(String dossierId, String fileId) {
InputStreamResource inputStreamResource;
@@ -73,6 +82,13 @@ public class RedactionStorageService {
return null;
}
return deserializeImportedRedactions(inputStreamResource);
}
@Timed("redactmanager_deserializeImportedRedactions")
private ImportedRedactions deserializeImportedRedactions(InputStreamResource inputStreamResource) {
try {
return dslJson.deserialize(ImportedRedactions.class, inputStreamResource.getInputStream());
} catch (IOException e) {
@@ -81,6 +97,7 @@ public class RedactionStorageService {
}
@Timed("redactmanager_getRedactionLog")
public RedactionLog getRedactionLog(String dossierId, String fileId) {
InputStreamResource inputStreamResource;
@@ -91,6 +108,13 @@ public class RedactionStorageService {
return null;
}
return deserializeRedactionLog(inputStreamResource);
}
@Timed("redactmanager_deserializeRedactionLog")
private RedactionLog deserializeRedactionLog(InputStreamResource inputStreamResource) {
try {
return dslJson.deserialize(RedactionLog.class, inputStreamResource.getInputStream());
} catch (IOException e) {
@@ -99,6 +123,7 @@ public class RedactionStorageService {
}
@Timed("redactmanager_getText")
public Text getText(String dossierId, String fileId) {
InputStreamResource inputStreamResource;
@@ -109,6 +134,13 @@ public class RedactionStorageService {
return null;
}
return deserializeText(inputStreamResource);
}
@Timed("redactmanager_deserializeText")
private Text deserializeText(InputStreamResource inputStreamResource) {
try {
return dslJson.deserialize(Text.class, inputStreamResource.getInputStream());
} catch (IOException e) {
@@ -117,6 +149,7 @@ public class RedactionStorageService {
}
@Timed("redactmanager_getNerEntities")
public NerEntities getNerEntities(String dossierId, String fileId) {
InputStreamResource inputStreamResource;
@@ -126,6 +159,13 @@ public class RedactionStorageService {
throw new NotFoundException("NER Entities are not available.");
}
return deserializeNerEntities(inputStreamResource);
}
@Timed("redactmanager_deserializeNerEntities")
private NerEntities deserializeNerEntities(InputStreamResource inputStreamResource) {
try {
return dslJson.deserialize(NerEntities.class, inputStreamResource.getInputStream());
} catch (IOException e) {
@@ -134,15 +174,27 @@ public class RedactionStorageService {
}
@Timed("redactmanager_getSectionGrid")
public SectionGrid getSectionGrid(String dossierId, String fileId) {
InputStreamResource inputStreamResource;
try {
var sectionGrid = storageService.getObject(StorageIdUtils.getStorageId(dossierId, fileId, FileType.SECTION_GRID));
return dslJson.deserialize(SectionGrid.class, sectionGrid.getInputStream());
inputStreamResource = storageService.getObject(StorageIdUtils.getStorageId(dossierId, fileId, FileType.SECTION_GRID));
} catch (StorageObjectDoesNotExist e) {
throw new NotFoundException("Section Grid is not available.");
}
return deserializeSectionGrid(inputStreamResource);
}
@Timed("redactmanager_deserializeSectionGrid")
private SectionGrid deserializeSectionGrid(InputStreamResource inputStreamResource) {
try {
return dslJson.deserialize(SectionGrid.class, inputStreamResource.getInputStream());
} catch (IOException e) {
throw new RuntimeException("Could not convert RedactionLog", e);
throw new RuntimeException("Could not convert SectionGrid", e);
}
}
@@ -1,6 +1,7 @@
package com.iqser.red.service.redaction.v1.server;
import com.amazonaws.services.s3.AmazonS3;
import com.fasterxml.jackson.core.type.TypeReference;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.AnnotationStatus;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.Comment;
@@ -29,7 +30,9 @@ import com.iqser.red.service.redaction.v1.server.redaction.utils.TextNormalizati
import com.iqser.red.service.redaction.v1.server.storage.RedactionStorageService;
import com.iqser.red.storage.commons.StorageAutoConfiguration;
import com.iqser.red.storage.commons.service.StorageService;
import lombok.SneakyThrows;
import org.apache.commons.io.IOUtils;
import org.junit.After;
import org.junit.Before;
@@ -58,6 +61,7 @@ import java.io.*;
import java.net.URL;
import java.nio.charset.StandardCharsets;
import java.time.OffsetDateTime;
import java.time.ZoneOffset;
import java.util.*;
import java.util.stream.Collectors;
@@ -214,16 +218,16 @@ public class RedactionIntegrationTest {
}
private void mockDictionaryCalls(Long version){
private void mockDictionaryCalls(Long version) {
when(dictionaryClient.getDictionaryForType(VERTEBRATE + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(VERTEBRATE, false));
when(dictionaryClient.getDictionaryForType(ADDRESS + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(ADDRESS, false));
when(dictionaryClient.getDictionaryForType(AUTHOR + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(AUTHOR, false));
when(dictionaryClient.getDictionaryForType(SPONSOR + ":" + TEST_DOSSIER_TEMPLATE_ID,version)).thenReturn(getDictionaryResponse(SPONSOR, false));
when(dictionaryClient.getDictionaryForType(SPONSOR + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(SPONSOR, false));
when(dictionaryClient.getDictionaryForType(NO_REDACTION_INDICATOR + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(NO_REDACTION_INDICATOR, false));
when(dictionaryClient.getDictionaryForType(REDACTION_INDICATOR + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(REDACTION_INDICATOR, false));
when(dictionaryClient.getDictionaryForType(HINT_ONLY + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(HINT_ONLY, false));
when(dictionaryClient.getDictionaryForType(MUST_REDACT + ":" + TEST_DOSSIER_TEMPLATE_ID,version)).thenReturn(getDictionaryResponse(MUST_REDACT, false));
when(dictionaryClient.getDictionaryForType(MUST_REDACT + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(MUST_REDACT, false));
when(dictionaryClient.getDictionaryForType(PUBLISHED_INFORMATION + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(PUBLISHED_INFORMATION, false));
when(dictionaryClient.getDictionaryForType(TEST_METHOD + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(TEST_METHOD, false));
when(dictionaryClient.getDictionaryForType(PII + ":" + TEST_DOSSIER_TEMPLATE_ID, version)).thenReturn(getDictionaryResponse(PII, false));
@@ -238,6 +242,7 @@ public class RedactionIntegrationTest {
}
@Test
public void test270Rotated() {
@@ -491,14 +496,12 @@ public class RedactionIntegrationTest {
deleted.add("David Chubb");
deleted.add("mouse");
reanlysisVersions.put("mouse", 3L);
when(dictionaryClient.getVersion(TEST_DOSSIER_TEMPLATE_ID)).thenReturn(3L);
when(dictionaryClient.getDictionaryForType(VERTEBRATE, null)).thenReturn(getDictionaryResponse(VERTEBRATE, false));
start = System.currentTimeMillis();
ManualRedactions manualRedactions = new ManualRedactions();
@@ -669,7 +672,6 @@ public class RedactionIntegrationTest {
when(dictionaryClient.getDictionaryForType(VERTEBRATE, null)).thenReturn(getDictionaryResponse(VERTEBRATE, false));
start = System.currentTimeMillis();
ManualRedactions manualRedactions = new ManualRedactions();
@@ -715,6 +717,71 @@ public class RedactionIntegrationTest {
}
@Test
public void testRemovePublishedInformations() throws IOException {
long start = System.currentTimeMillis();
ClassPathResource colorsResource = new ClassPathResource("colors/colors.json");
var colors = objectMapper.readValue(colorsResource.getInputStream(), Colors.class);
ClassPathResource typeResource = new ClassPathResource("colors/types.json");
TypeReference<List<Type>> typeRefForTypes = new TypeReference<>() {
};
List<Type> types = objectMapper.readValue(typeResource.getInputStream(), typeRefForTypes);
AnalyzeRequest request = prepareStorage("files/new/PublishedInformationTest.pdf");
analyzeService.analyzeDocumentStructure(new StructureAnalyzeRequest(request.getDossierId(), request.getFileId()));
ManualRedactions manualRedactions = new ManualRedactions();
manualRedactions.getIdsToRemove()
.add(IdRemoval.builder()
.annotationId("308dab9015bfafd911568cffe0a7f7de")
.fileId(TEST_FILE_ID)
.status(AnnotationStatus.APPROVED)
.requestDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 07, 475479, ZoneOffset.UTC))
.processedDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 07, 483651, ZoneOffset.UTC))
.build());
manualRedactions.getForceRedactions()
.add(ManualForceRedaction.builder()
.annotationId("0b56ea1a87c83f351df177315af94f0d")
.fileId(TEST_FILE_ID)
.status(AnnotationStatus.APPROVED)
.requestDate(OffsetDateTime.of(2022, 05, 23, 9, 30, 15, 4653, ZoneOffset.UTC))
.processedDate(OffsetDateTime.of(2022, 05, 23, 9, 30, 15, 794, ZoneOffset.UTC))
.build());
manualRedactions.getIdsToRemove()
.add(IdRemoval.builder()
.annotationId("0b56ea1a87c83f351df177315af94f0d")
.fileId(TEST_FILE_ID)
.status(AnnotationStatus.APPROVED)
.requestDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 23, 961721, ZoneOffset.UTC))
.processedDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 23, 96528, ZoneOffset.UTC))
.build());
request.setManualRedactions(manualRedactions);
AnalyzeResult result = analyzeService.analyze(request);
AnnotateResponse annotateResponse = annotationService.annotate(AnnotateRequest.builder()
.manualRedactions(manualRedactions)
.colors(colors)
.types(types)
.dossierId(TEST_DOSSIER_ID)
.fileId(TEST_FILE_ID)
.build());
try (FileOutputStream fileOutputStream = new FileOutputStream(OsUtils.getTemporaryDirectory() + "/Annotated.pdf")) {
fileOutputStream.write(annotateResponse.getDocument());
}
long end = System.currentTimeMillis();
System.out.println("duration: " + (end - start));
System.out.println("numberOfPages: " + result.getNumberOfPages());
}
@Test
public void testTableRedaction() throws IOException {
@@ -740,6 +807,89 @@ public class RedactionIntegrationTest {
}
@Test
public void testRotations() throws IOException {
System.out.println("testTableRedaction");
long start = System.currentTimeMillis();
AnalyzeRequest request = prepareStorage("files/new/RotateTestFile.pdf");
analyzeService.analyzeDocumentStructure(new StructureAnalyzeRequest(request.getDossierId(), request.getFileId()));
AnalyzeResult result = analyzeService.analyze(request);
AnnotateResponse annotateResponse = annotationService.annotate(AnnotateRequest.builder()
.dossierId(TEST_DOSSIER_ID)
.fileId(TEST_FILE_ID)
.build());
try (FileOutputStream fileOutputStream = new FileOutputStream(OsUtils.getTemporaryDirectory() + "/Annotated.pdf")) {
fileOutputStream.write(annotateResponse.getDocument());
}
long end = System.currentTimeMillis();
System.out.println("duration: " + (end - start));
System.out.println("numberOfPages: " + result.getNumberOfPages());
}
@Test
public void testFindDictionaryEntryInResizedEntryPosition() throws IOException {
System.out.println("testResizeDictFound");
ClassPathResource colorsResource = new ClassPathResource("colors/colors.json");
var colors = objectMapper.readValue(colorsResource.getInputStream(), Colors.class);
ClassPathResource typeResource = new ClassPathResource("colors/types.json");
TypeReference<List<Type>> typeRefForTypes = new TypeReference<>() {
};
List<Type> types = objectMapper.readValue(typeResource.getInputStream(), typeRefForTypes);
long start = System.currentTimeMillis();
AnalyzeRequest request = prepareStorage("files/new/S157.pdf");
ManualRedactions manualRedactions = new ManualRedactions();
manualRedactions.getResizeRedactions()
.add(ManualResizeRedaction.builder()
.annotationId("ca2b437e2480a4b5966cb8386020d454")
.fileId(TEST_FILE_ID)
.status(AnnotationStatus.APPROVED)
.requestDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 07, 475479, ZoneOffset.UTC))
.processedDate(OffsetDateTime.of(2022, 05, 23, 8, 30, 07, 483651, ZoneOffset.UTC))
.value("Bera P")
.positions(List.of(new Rectangle(382.12485f, 235.94246f, 8.768621f, -21.632504f, 1)))
.textBefore("")
.textAfter("")
.build());
request.setManualRedactions(manualRedactions);
analyzeService.analyzeDocumentStructure(new StructureAnalyzeRequest(request.getDossierId(), request.getFileId()));
AnalyzeResult result = analyzeService.analyze(request);
var redactionLog = redactionStorageService.getRedactionLog(TEST_DOSSIER_ID, TEST_FILE_ID);
AnnotateResponse annotateResponse = annotationService.annotate(AnnotateRequest.builder()
.manualRedactions(manualRedactions)
.colors(colors)
.types(types)
.dossierId(TEST_DOSSIER_ID)
.fileId(TEST_FILE_ID)
.build());
try (FileOutputStream fileOutputStream = new FileOutputStream(OsUtils.getTemporaryDirectory() + "/Annotated.pdf")) {
fileOutputStream.write(annotateResponse.getDocument());
}
long end = System.currentTimeMillis();
System.out.println("duration: " + (end - start));
System.out.println("numberOfPages: " + result.getNumberOfPages());
}
@Test
public void testManualRedaction() throws IOException {
@@ -1361,7 +1511,7 @@ public class RedactionIntegrationTest {
typeColorMap.put(LOGO, "#ffe187");
typeColorMap.put(FORMULA, "#ffe187");
typeColorMap.put(SIGNATURE, "#ffe187");
typeColorMap.put(IMPORTED_REDACTION, "#32a852");
typeColorMap.put(IMPORTED_REDACTION, "#fcfbe6");
hintTypeMap.put(VERTEBRATE, true);
hintTypeMap.put(ADDRESS, false);
@@ -1518,8 +1668,8 @@ public class RedactionIntegrationTest {
public void testImportedRedactions() throws IOException {
String outputFileName = OsUtils.getTemporaryDirectory() + "/Annotated.pdf";
ClassPathResource pdfFileResource = new ClassPathResource("files/ImportedRedactions/ImportedRedactions.pdf");
ClassPathResource importedRedactions = new ClassPathResource("files/ImportedRedactions/ImportedRedactions.json");
ClassPathResource pdfFileResource = new ClassPathResource("files/ImportedRedactions/RotateTestFile_without_highlights.pdf");
ClassPathResource importedRedactions = new ClassPathResource("files/ImportedRedactions/RotateTestFile_without_highlights.IMPORTED_REDACTIONS.json");
AnalyzeRequest request = prepareStorage(pdfFileResource.getInputStream());
storageService.storeObject(RedactionStorageService.StorageIdUtils.getStorageId(TEST_DOSSIER_ID, TEST_FILE_ID, FileType.IMPORTED_REDACTIONS), IOUtils.toByteArray(importedRedactions.getInputStream()));
@@ -1534,20 +1684,33 @@ public class RedactionIntegrationTest {
.fileId(TEST_FILE_ID)
.build());
redactionLog.getRedactionLogEntry().forEach(entry -> {
if (entry.getValue() == null) {
return;
}
if (entry.getValue().equals("David")){
assertThat(entry.getImportedRedactionIntersections().size()).isEqualTo(1);
}
if (entry.getValue().equals("annotation")){
assertThat(entry.getImportedRedactionIntersections().size()).isEqualTo(0);
}
});
try (FileOutputStream fileOutputStream = new FileOutputStream(outputFileName)) {
fileOutputStream.write(annotateResponse.getDocument());
}
}
@Test
public void testExpandByPrefixRegEx() throws IOException {
assertThat(dictionary.get(AUTHOR).contains("Robinson"));
assertThat(! dictionary.get(AUTHOR).contains("Mrs. Robinson"));
assertThat(!dictionary.get(AUTHOR).contains("Mrs. Robinson"));
assertThat(dictionary.get(AUTHOR).contains("Bojangles"));
assertThat(! dictionary.get(AUTHOR).contains("Mr. Bojangles"));
assertThat(!dictionary.get(AUTHOR).contains("Mr. Bojangles"));
assertThat(dictionary.get(AUTHOR).contains("Tambourine Man"));
assertThat(! dictionary.get(AUTHOR).contains("Mr. Tambourine Man"));
assertThat(!dictionary.get(AUTHOR).contains("Mr. Tambourine Man"));
String fileName = "files/mr-mrs.pdf";
String outputFileName = OsUtils.getTemporaryDirectory() + "/Annotated.pdf";
@@ -1578,6 +1741,7 @@ public class RedactionIntegrationTest {
assertThat(values).contains("Mr. Tambourine Man");
}
@SneakyThrows
private AnalyzeRequest prepareStorage(InputStream stream) {
@@ -1,6 +1,10 @@
package com.iqser.red.service.redaction.v1.server.annotate;
import java.util.List;
import com.iqser.red.service.persistence.service.v1.api.model.annotations.ManualRedactions;
import com.iqser.red.service.persistence.service.v1.api.model.dossiertemplate.configuration.Colors;
import com.iqser.red.service.persistence.service.v1.api.model.dossiertemplate.type.Type;
import lombok.AllArgsConstructor;
import lombok.Builder;
@@ -17,4 +21,6 @@ public class AnnotateRequest {
private String dossierTemplateId;
private String fileId;
private ManualRedactions manualRedactions;
private Colors colors;
private List<Type> types;
}
@@ -46,6 +46,8 @@ public class AnnotationService {
.manualRedactions(annotateRequest.getManualRedactions())
.dossierId(annotateRequest.getDossierId())
.dossierTemplateId(annotateRequest.getDossierTemplateId())
.colors(annotateRequest.getColors())
.types(annotateRequest.getTypes())
.build());
var sectionsGrid = redactionStorageService.getSectionGrid(annotateRequest.getDossierId(), annotateRequest.getFileId());
@@ -116,7 +118,7 @@ public class AnnotationService {
annotation.setRectangle(pdRectangle);
annotation.setQuadPoints(toQuadPoints(rectangles, mediaBox, cropBox));
if (!redactionLogEntry.isHint()) {
annotation.setContents(createAnnotationContent(redactionLogEntry));
annotation.setContents(redactionLogEntry.getValue() + " " +createAnnotationContent(redactionLogEntry));
}
annotation.setTitlePopup(redactionLogEntry.getId());
annotation.setAnnotationName(redactionLogEntry.getId());
@@ -96,7 +96,7 @@ public class PdfSegmentationServiceTest {
Map<Integer, List<PdfImage>> images = new HashMap<>();
imageServiceResponse.getData().stream().forEach(imageMetadata -> {
images.computeIfAbsent(imageMetadata.getPosition().getPageNumber() ,x -> new ArrayList<>())
.add(new PdfImage(new RedRectangle2D(imageMetadata.getPosition().getX1(), imageMetadata.getPosition().getY1(), imageMetadata.getGeometry().getWidth(), imageMetadata.getGeometry().getHeight()), ImageType.valueOf(imageMetadata.getClassification().getLabel().toUpperCase(Locale.ROOT)), imageMetadata.getPosition().getPageNumber()));
.add(new PdfImage(new RedRectangle2D(imageMetadata.getPosition().getX1(), imageMetadata.getPosition().getY1(), imageMetadata.getGeometry().getWidth(), imageMetadata.getGeometry().getHeight()), ImageType.valueOf(imageMetadata.getClassification().getLabel().toUpperCase(Locale.ROOT)), imageMetadata.isAlpha(), imageMetadata.getPosition().getPageNumber()));
});
System.out.println("object");
@@ -0,0 +1,85 @@
package com.iqser.red.service.redaction.v1.server.stringmatching;
import lombok.AllArgsConstructor;
import lombok.EqualsAndHashCode;
import lombok.SneakyThrows;
import org.ahocorasick.trie.Trie;
import org.apache.commons.io.IOUtils;
import org.junit.Test;
import org.junit.runner.RunWith;
import org.springframework.core.io.ClassPathResource;
import org.springframework.test.context.junit4.SpringRunner;
import java.util.HashSet;
import java.util.Set;
import java.util.regex.Pattern;
import java.util.stream.Collectors;
import static org.assertj.core.api.AssertionsForClassTypes.assertThat;
@RunWith(SpringRunner.class)
public class StringMatchingPerformanceTest {
@Test
@SneakyThrows
public void testStringPerformance() {
String text = IOUtils.toString(new ClassPathResource("stringmatching/hamlet.txt").getInputStream()).toLowerCase();
Set<String> dictionary = IOUtils.readLines(new ClassPathResource("stringmatching/names.txt").getInputStream())
.stream()
.map(String::toLowerCase).collect(Collectors.toSet());
System.out.println("Loaded text has a length of " + text.length() + " symbols");
System.out.println("Dictionary has " + dictionary.size() + " entries");
var patterns = dictionary.stream()
.map(p -> Pattern.compile(Pattern.quote(p))).collect(Collectors.toList());
var trie = Trie.builder().ignoreCase().addKeywords(dictionary).build();
// 1. Naive approach
long t1 = System.currentTimeMillis();
var naiveIndexes = new HashSet<Index>();
for (var entry : dictionary) {
var startIndex = 0;
do {
startIndex = text.indexOf(entry, startIndex + 1);
if (startIndex != -1) {
naiveIndexes.add(new Index(startIndex, startIndex + entry.length()));
}
} while (startIndex != -1);
}
long t2 = System.currentTimeMillis();
System.out.println("Naive approach found " + naiveIndexes.size() + " entries in " + (t2 - t1) + "ms");
// 2. Boyer Moore
t1 = System.currentTimeMillis();
var boyerMooreIndexes = new HashSet<Index>();
for (var pattern : patterns) {
boyerMooreIndexes.addAll(pattern.matcher(text).results().map(r -> new Index(r.start(), r.end())).collect(Collectors.toList()));
}
t2 = System.currentTimeMillis();
System.out.println("Boyer Moore found " + boyerMooreIndexes.size() + " entries in " + (t2 - t1) + "ms");
// 3. Aho Corasick
t1 = System.currentTimeMillis();
var result = trie.parseText(text);
var ahoCorasickIndexes = result.stream().map(r -> new Index(r.getStart(), r.getEnd() + 1)).collect(Collectors.toSet());
t2 = System.currentTimeMillis();
System.out.println("Aho Corasick found " + ahoCorasickIndexes.size() + " entries in " + (t2 - t1) + "ms");
// Assert that all algorithms are equal
assertThat(naiveIndexes).isEqualTo(boyerMooreIndexes).isEqualTo(ahoCorasickIndexes);
}
@AllArgsConstructor
@EqualsAndHashCode(of = {"start", "end"})
public static class Index {
int start;
int end;
}
}
@@ -0,0 +1,13 @@
{
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"defaultColor": "#9398a0",
"requestAdd": "#04b093",
"requestRemove": "#04b093",
"notRedacted": "#c498fa",
"analysisColor": "#dd4d50",
"updatedColor": "#fdbd00",
"dictionaryRequestColor": "#5b97db",
"manualRedactionColor": "#9398a0",
"previewColor": "#9398a0",
"ignoredHintColor": "#e7d4ff"
}
@@ -0,0 +1,223 @@
[
{
"id": "CBI_address:31039447-9040-4376-9ca7-614e56b284b9",
"type": "CBI_address",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 140,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "All site names and addresses, and location (e.g. Syngenta, Monthey, GPS Co-ordinates, Mr Smith of … providing the…). Except addresses in published literature and the applicant address.",
"addToDictionaryAction": true,
"label": "CBI Address",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "CBI_author:31039447-9040-4376-9ca7-614e56b284b9",
"type": "CBI_author",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 130,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "All authors named in the study documentation. Except names in published literature.",
"addToDictionaryAction": true,
"label": "CBI Author",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "PII:31039447-9040-4376-9ca7-614e56b284b9",
"type": "PII",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 150,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "Not authors but listed in the document: Names, signatures, telephone, email etc.; e.g. Reg Manager, QA Manager",
"addToDictionaryAction": true,
"label": "PII",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "formula:31039447-9040-4376-9ca7-614e56b284b9",
"type": "formula",
"hexColor": "#036ffc",
"recommendationHexColor": "#8df06c",
"rank": 1002,
"isHint": true,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Empty dictionary used to configure formula colors.",
"addToDictionaryAction": false,
"label": "Formula",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": false
},
{
"id": "isHint_only:31039447-9040-4376-9ca7-614e56b284b9",
"type": "isHint_only",
"hexColor": "#fa98f7",
"recommendationHexColor": "#8df06c",
"rank": 50,
"isHint": true,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Entries of this dictionary will be highlighted only",
"addToDictionaryAction": false,
"label": "isHint Only",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "image:31039447-9040-4376-9ca7-614e56b284b9",
"type": "image",
"hexColor": "#bdd6ff",
"recommendationHexColor": "#8df06c",
"rank": 999,
"isHint": true,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Empty dictionary used to configure image colors.",
"addToDictionaryAction": false,
"label": "Image",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": false
},
{
"id": "logo:31039447-9040-4376-9ca7-614e56b284b9",
"type": "logo",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 1001,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Empty dictionary used to configure logo colors.",
"addToDictionaryAction": false,
"label": "Logo",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": false
},
{
"id": "must_redact:31039447-9040-4376-9ca7-614e56b284b9",
"type": "must_redact",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 100,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Entries of this dictionary get redacted wherever found.",
"addToDictionaryAction": false,
"label": "Must Redact",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "ocr:31039447-9040-4376-9ca7-614e56b284b9",
"type": "ocr",
"hexColor": "#bdd6ff",
"recommendationHexColor": "#8df06c",
"rank": 1000,
"isHint": true,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Empty dictionary used to configure ocr colors.",
"addToDictionaryAction": false,
"label": "Ocr",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": false
},
{
"id": "signature:31039447-9040-4376-9ca7-614e56b284b9",
"type": "signature",
"hexColor": "#9398a0",
"recommendationHexColor": "#8df06c",
"rank": 1003,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": true,
"isRecommendation": false,
"description": "Empty dictionary used to configure signature colors.",
"addToDictionaryAction": false,
"label": "Signature",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": false
},
{
"id": "imported_redaction:31039447-9040-4376-9ca7-614e56b284b9",
"type": "imported_redaction",
"hexColor": "#f0f0c0",
"recommendationHexColor": "#8df06c",
"rank": 9999,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "Redaction Annotations that were imported from documents",
"addToDictionaryAction": false,
"label": "Imported Redaction",
"hasDictionary": false,
"systemManaged": true,
"autoHideSkipped": true
},
{
"id": "published_information:31039447-9040-4376-9ca7-614e56b284b9",
"type": "published_information",
"hexColor": "#85ebff",
"recommendationHexColor": "#8df06c",
"rank": 70,
"isHint": true,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "Manual managed list of public journals and papers that need no redaction",
"addToDictionaryAction": true,
"label": "Published Information",
"hasDictionary": true,
"systemManaged": false,
"autoHideSkipped": false
},
{
"id": "dossier_redaction:31039447-9040-4376-9ca7-614e56b284b9:5dfb2724-74a4-4a1a-a1eb-165e7943ffcd",
"type": "dossier_redaction",
"hexColor": "#9398a0",
"recommendationHexColor": null,
"rank": 1500,
"isHint": false,
"dossierTemplateId": "31039447-9040-4376-9ca7-614e56b284b9",
"isCaseInsensitive": false,
"isRecommendation": false,
"description": "Entries in this dictionary will only be redacted in this dossier",
"addToDictionaryAction": false,
"label": "Dossier Redaction",
"hasDictionary": true,
"systemManaged": true,
"autoHideSkipped": false
}
]
@@ -9,3 +9,7 @@ Naka-27 Aomachi, Nomi, Ishikawa 923-1101, Japan, JP
Özgür U. Reyhan
Sude Halide Nurullah
Xinyi Y. Tao
Dorn
Prasher
David
annotation
@@ -86,3 +86,4 @@ Toxicol Sci
Toxicol Sci.
Toxicol Sci. 1
Test Ignored Hint Published Information
Workshop
@@ -1 +0,0 @@
{"importedRedactions":{"1":[{"id":"755273b5ddc6933eff7f6b6e055c803f","positions":[{"topLeft":{"x":275.0,"y":520.504},"width":55.675,"height":15.582,"page":1}]},{"id":"5e517d2c890272fc140578a26971b459","positions":[{"topLeft":{"x":164.0,"y":316.492},"width":50.7,"height":15.658,"page":1},{"topLeft":{"x":127.0,"y":302.613},"width":119.675,"height":14.902,"page":1},{"topLeft":{"x":81.0,"y":289.622},"width":210.025,"height":14.846,"page":1},{"topLeft":{"x":136.0,"y":275.912},"width":94.2,"height":13.033,"page":1},{"topLeft":{"x":158.0,"y":261.402},"width":52.793,"height":16.218,"page":1},{"topLeft":{"x":177.0,"y":247.413},"width":15.507,"height":16.148,"page":1}]}]}}
@@ -0,0 +1,236 @@
{
"importedRedactions": {
"1": [
{
"id": "e9b15eafe60232957c4805439d41b4bb",
"positions": [
{
"topLeft": {
"x": 173.4,
"y": 396.0
},
"width": -29.5,
"height": -8.1,
"page": 1
}
]
},
{
"id": "5dc06b8f402a5ea6f4a71e7f0ae5f42e",
"positions": [
{
"topLeft": {
"x": 189.4,
"y": 694.0
},
"width": 28.0,
"height": -12.1,
"page": 1
}
]
},
{
"id": "42ced0b82072cd154f05f917b5063c23",
"positions": [
{
"topLeft": {
"x": 181.8,
"y": 593.3
},
"width": -10.1,
"height": -30.0,
"page": 1
}
]
},
{
"id": "8569561e9cca5d96f2fe966d1f470fbe",
"positions": [
{
"topLeft": {
"x": 168.0,
"y": 180.1
},
"width": 10.1,
"height": 30.8,
"page": 1
}
]
}
],
"2": [
{
"id": "22f46754f11d153e4693c64dd95a0f5f",
"positions": [
{
"topLeft": {
"x": 398.8,
"y": 302.1
},
"width": -27.6,
"height": -12.1,
"page": 2
}
]
},
{
"id": "27386f133b866d27a25ceafcb4bd4774",
"positions": [
{
"topLeft": {
"x": 153.1,
"y": 258.2
},
"width": 25.5,
"height": -10.1,
"page": 2
}
]
},
{
"id": "294b8d00efd202e4aa51f6fc5e225e80",
"positions": [
{
"topLeft": {
"x": 46.3,
"y": 304.8
},
"width": 10.1,
"height": -25.5,
"page": 2
}
]
},
{
"id": "2f7142be7b26f9efd1cd608ceaae8f25",
"positions": [
{
"topLeft": {
"x": 523.6,
"y": 238.9
},
"width": -10.1,
"height": 25.5,
"page": 2
}
]
}
],
"3": [
{
"id": "2566b5f6f11eb7609d0d9722513d6bec",
"positions": [
{
"topLeft": {
"x": 444.1,
"y": 415.4
},
"width": 28.0,
"height": -12.1,
"page": 3
}
]
},
{
"id": "84350bb01e0e93be7146a09f975c3065",
"positions": [
{
"topLeft": {
"x": 448.6,
"y": 615.9
},
"width": -10.1,
"height": -27.5,
"page": 3
}
]
},
{
"id": "e1dc39e4f5006fa11e6dde38003d4b08",
"positions": [
{
"topLeft": {
"x": 434.8,
"y": 208.2
},
"width": 10.1,
"height": 27.8,
"page": 3
}
]
},
{
"id": "75e67d849627e16becb66ec1ba51c7c8",
"positions": [
{
"topLeft": {
"x": 428.2,
"y": 117.5
},
"width": -29.4,
"height": -8.1,
"page": 3
}
]
}
],
"4": [
{
"id": "1c1a66ed056257e6e5ebc58b882acaee",
"positions": [
{
"topLeft": {
"x": 526.6,
"y": 487.8
},
"width": 10.1,
"height": 27.9,
"page": 4
}
]
},
{
"id": "7f79cfc276f448f5ee46c3fa2c84ed8b",
"positions": [
{
"topLeft": {
"x": 212.6,
"y": 504.9
},
"width": -27.5,
"height": -8.1,
"page": 4
}
]
},
{
"id": "64b11b7d6143f1deb4f1fd3bd72c0464",
"positions": [
{
"topLeft": {
"x": 69.5,
"y": 558.5
},
"width": -10.1,
"height": -27.9,
"page": 4
}
]
},
{
"id": "d77ea54e7d2b2276efa7cd9f1d08b411",
"positions": [
{
"topLeft": {
"x": 404.4,
"y": 545.8
},
"width": 27.1,
"height": -9.1,
"page": 4
}
]
}
]
}
}