Files
DocumentService/REQUIRED_FEATURES.md
TekH 0e88b349d7 Rebrand project: DocumentOperator to DocumentService
This commit implements a complete rebranding of the project:
- Updated all namespaces from `DocumentOperator` to `DocumentService`.
- Renamed file paths, embedded resources, and test data references.
- Updated configuration keys, logging paths, and Redis instance names.
- Revised documentation to reflect the new project name.
- Modified project and solution files to align with the new structure.
- Updated class names, DTOs, commands, queries, and handlers.
- Adjusted middleware, controllers, and API endpoints.
- Updated Swagger metadata and API titles to `DocumentService API`.
- Refactored test namespaces, resource paths, and embedded resources.
- Updated build and deployment configurations for the new name.
- Replaced all references to `DocumentOperator` in comments and literals.

These changes ensure consistency across the codebase and documentation.
2026-07-30 14:02:56 +02:00

11 KiB

DocumentService - Required Functions and Features

Project: DocumentService (DOC)
Ticket: DOC-1 - GDPicture and Nutrient Replacing
Owner: Hakan Tek
Last Updated: 03.07.2026


Overview

This document describes in detail all PDF processing functions that the DocumentService service must implement. These functions will be implemented using the DevExpress Office File API as a replacement for the GDPicture and Nutrient libraries.


1. Input Processing

1.1 PDF/A Validation

Purpose: Check whether a PDF file conforms to the PDF/A standard.

What is PDF/A?

  • ISO standard PDF format for long-term archiving
  • Versions: PDF/A-1, PDF/A-2, PDF/A-3
  • Has restrictions to guarantee future accessibility

Criteria to Check:

  • ✓ Are all fonts embedded? (required in PDF/A)
  • ✓ Is encryption present? (prohibited in PDF/A)
  • ✓ Are there external content links? (prohibited in PDF/A)
  • ✓ Does it contain JavaScript? (prohibited in PDF/A)
  • ✓ Is metadata correctly defined?
  • ✓ Are color profiles defined?
  • ✓ Is transparency usage compliant with the standard?

Suggested API Endpoint:

POST /api/pdf/validate-pdfa
Body: { "filePath": "string" } or multipart file upload
Response: {
  "isValid": true/false,
  "pdfaVersion": "PDF/A-2b",
  "errors": [],
  "warnings": []
}

Example Usage Scenario:

Input: invoice_2024.pdf
Output: {
  "isValid": true,
  "pdfaVersion": "PDF/A-3b",
  "pageCount": 5,
  "errors": [],
  "warnings": ["Document contains optional content"]
}

1.2 Attachment Check

Purpose: Detect whether embedded files (attachments) exist inside a PDF file.

What is an Attachment?

  • External files embedded inside a PDF
  • Examples: XML, Excel, images, other PDFs
  • Special use: XML data in ZUGFeRD/XRechnung e-invoices

PDF/A Versions and Attachments:

  • PDF/A-1: NO attachment support
  • PDF/A-2: NO attachment support
  • PDF/A-3: Attachment support AVAILABLE (most common use)

Suggested API Endpoint:

POST /api/pdf/has-attachments
Response: {
  "hasAttachments": true/false,
  "attachmentCount": 2,
  "attachments": [
    {
      "fileName": "factur-x.xml",
      "mimeType": "application/xml",
      "size": 12345,
      "description": "ZUGFeRD Invoice Data"
    }
  ]
}

1.3 Extract Attachments

Purpose: Extract all embedded files from a PDF and save them to the specified location.

Usage Scenarios:

  • Extracting XML from ZUGFeRD e-invoices
  • Retrieving attached documents from PDF/A-3 files
  • Extracting embedded data for automated processing

Suggested API Endpoint:

POST /api/pdf/extract-attachments
Body: {
  "filePath": "string",
  "outputPath": "string",
  "fileNamePattern": "{originalName}" // or "{index}_{originalName}"
}
Response: {
  "success": true,
  "extractedFiles": [
    {
      "fileName": "factur-x.xml",
      "savedPath": "C:\\Temp\\Attachments\\factur-x.xml",
      "size": 12345
    }
  ]
}

DevExpress Implementation Example:

using DevExpress.Pdf;

using (PdfDocumentProcessor processor = new PdfDocumentProcessor())
{
    processor.LoadDocument("invoice.pdf");
    
    foreach (PdfEmbeddedFile embeddedFile in processor.Document.EmbeddedFiles)
    {
        byte[] fileData = embeddedFile.GetData();
        string outputPath = Path.Combine(targetFolder, embeddedFile.FileName);
        File.WriteAllBytes(outputPath, fileData);
    }
}

1.4 PDF File Validity Check

Purpose: Check whether a file is a genuine PDF and verify its structural integrity.

Items to Check:

  • ✓ Is the file actually a PDF? (Magic bytes: %PDF-)
  • ✓ Is the PDF header valid?
  • ✓ Is the PDF structure not broken/corrupt?
  • ✓ Can the file be opened?
  • ✓ What is the PDF version? (1.4, 1.7, 2.0, etc.)
  • ✓ Page count
  • ✓ Can basic metadata be read?

Suggested API Endpoint:

POST /api/pdf/validate
Response: {
  "isValid": true,
  "pdfVersion": "1.7",
  "pageCount": 5,
  "fileSize": 524288,
  "encrypted": false,
  "errors": [],
  "metadata": {
    "title": "Invoice 2024",
    "author": "Company XYZ",
    "creationDate": "2024-01-15T10:30:00"
  }
}

Error Scenarios:

Scenario 1 - Corrupt file:
{
  "isValid": false,
  "errors": ["PDF header is missing or corrupted"]
}

Scenario 2 - Wrong format:
{
  "isValid": false,
  "errors": ["File is not a PDF (detected: JPEG image)"]
}

2. PDF Processing Operations

2.1 PDF Merge (Concatenate/Merge)

Purpose: Merge multiple PDF files into a single PDF.

Usage Areas:

  • signFLOW (Envelope Generator): Merging multiple documents into a single envelope
  • ErgebnisberichtCreator: Merging report sections
  • ResultHandler (windream): Document merging

Suggested API Endpoint:

POST /api/pdf/merge
Body: {
  "sourceFiles": ["file1.pdf", "file2.pdf", "file3.pdf"],
  "outputPath": "merged.pdf",
  "options": {
    "addBookmarks": true,
    "preserveMetadata": true,
    "compressionLevel": "medium"
  }
}
Response: {
  "success": true,
  "outputPath": "merged.pdf",
  "pageCount": 15,
  "fileSize": 1048576
}

Features:

  • Preserve page order
  • Add bookmarks (for each source file)
  • Merge metadata
  • Preserve PDF/A compliance

2.2 Stamping

Purpose: Add a stamp to PDF pages - e.g. "APPROVED", "CONFIDENTIAL", "DRAFT"

Usage Area:

  • ErgebnisberichtCreator: Approval/status stamps on reports

Suggested API Endpoint:

POST /api/pdf/stamp
Body: {
  "filePath": "document.pdf",
  "stamp": {
    "text": "APPROVED",
    "position": "TopRight", // TopLeft, TopRight, BottomLeft, BottomRight, Center
    "pages": "all", // or "1,3,5" or "1-5"
    "color": "#FF0000",
    "opacity": 0.5,
    "rotation": 45,
    "fontSize": 48
  }
}

Stamp Types:

  • Text stamp
  • Image stamp (logo, signature)
  • QR code stamp
  • Date/time stamp

2.3 Annotation

Purpose: Add comments, highlights, notes, and drawings to a PDF.

Usage Area:

  • signFLOW: Signature fields, comments, highlights

Annotation Types:

  • Text annotations (comments)
  • Highlight
  • Underline
  • Strikeout
  • Shapes (rectangle, circle, arrow)
  • Stamps (predefined stamps)

Suggested API Endpoint:

POST /api/pdf/annotate
Body: {
  "filePath": "document.pdf",
  "annotations": [
    {
      "type": "highlight",
      "page": 1,
      "rect": {"x": 100, "y": 200, "width": 200, "height": 20},
      "color": "#FFFF00"
    },
    {
      "type": "text",
      "page": 1,
      "position": {"x": 100, "y": 250},
      "text": "This section is important!",
      "author": "Hakan Tek"
    }
  ]
}

2.4 PDF/A Conversion

Purpose: PDF ↔ PDF/A format conversion.

Usage Area:

  • taskFLOW: Converting PDF/A to standard PDF (optional)

Two-Way Operation:

A) PDF → PDF/A (For archiving)

POST /api/pdf/convert-to-pdfa
Body: {
  "filePath": "document.pdf",
  "pdfaLevel": "PDF/A-2b", // or PDF/A-1b, PDF/A-3b
  "embedFonts": true,
  "colorProfile": "sRGB"
}

B) PDF/A → PDF (Remove restrictions)

POST /api/pdf/convert-from-pdfa
Body: {
  "filePath": "document-pdfa.pdf",
  "removeRestrictions": true
}

During Conversion:

  • Embed all fonts
  • Remove JavaScript
  • Resolve external references
  • Add color profiles
  • Metadata standardization

3. PDF Rendering (Preview/Rendering)

3.1 PDF Preview

Purpose: Render PDF pages and return them as images (PNG/JPEG).

Usage Areas:

  • taskFLOW: PDF preview
  • fileFLOW: PDF preview
  • easyFLOW: PDF preview
  • orgFLOW: PDF preview

Suggested API Endpoint:

POST /api/pdf/render
Body: {
  "filePath": "document.pdf",
  "page": 1, // or "all"
  "format": "png", // or "jpeg"
  "dpi": 150,
  "width": 800, // optional, aspect ratio preserved
  "quality": 85 // for JPEG
}
Response: {
  "images": [
    {
      "page": 1,
      "base64": "iVBORw0KGgoAAAANS...",
      "width": 800,
      "height": 1132
    }
  ]
}

Features:

  • Page selection (single page or all)
  • Configurable DPI
  • Format selection (PNG, JPEG)
  • Thumbnail generation

4. Advanced Functions

4.1 Text Extraction

Purpose: Extract plain text from a PDF.

POST /api/pdf/extract-text
Response: {
  "text": "Extracted text content...",
  "pageTexts": [
    { "page": 1, "text": "Page 1 content..." },
    { "page": 2, "text": "Page 2 content..." }
  ]
}

4.2 Metadata Read/Write

Purpose: Read and update PDF metadata.

GET /api/pdf/metadata?filePath=document.pdf
Response: {
  "title": "Invoice 2024",
  "author": "Company XYZ",
  "subject": "Monthly Invoice",
  "keywords": ["invoice", "payment"],
  "creator": "Microsoft Word",
  "producer": "Adobe PDF Library",
  "creationDate": "2024-01-15T10:30:00",
  "modificationDate": "2024-01-15T14:45:00"
}

POST /api/pdf/metadata
Body: {
  "filePath": "document.pdf",
  "metadata": {
    "title": "Updated Title",
    "author": "New Author"
  }
}

4.3 Form Field Operations

Purpose: Read and fill fields in PDF forms.

GET /api/pdf/form-fields
Response: {
  "fields": [
    { "name": "customerName", "type": "text", "value": "" },
    { "name": "invoiceDate", "type": "text", "value": "" },
    { "name": "approved", "type": "checkbox", "value": false }
  ]
}

POST /api/pdf/fill-form
Body: {
  "filePath": "form.pdf",
  "fields": {
    "customerName": "ACME Corp",
    "invoiceDate": "2024-07-03",
    "approved": true
  },
  "flatten": true // Make fields non-editable
}

5. Priority Order

Phase 1 - Core Functions (PRIORITY)

  1. ✓ PDF Validation (1.4)
  2. ✓ PDF/A Validation (1.1)
  3. ✓ Attachment Check (1.2)
  4. ✓ Extract Attachments (1.3)
  5. ✓ PDF Merge (2.1)

Phase 2 - Advanced Functions

  1. ✓ Stamping (2.2)
  2. ✓ Annotation (2.3)
  3. ✓ PDF Preview/Rendering (3.1)

Phase 3 - Conversion and Extras

  1. ✓ PDF/A Conversion (2.4)
  2. ✓ Text Extraction (4.1)
  3. ✓ Metadata Operations (4.2)
  4. ✓ Form Field Operations (4.3)

6. Technical Requirements

6.1 Library to Use

  • DevExpress Office File API (PDF Document API)
  • NuGet: DevExpress.Pdf
  • License: Existing DevExpress license

6.2 REST API Requirements

  • HTTP REST endpoints
  • JSON request/response
  • Swagger/OpenAPI documentation
  • File upload support (multipart/form-data)

6.3 Performance Goals

  • Asynchronous processing (async/await)
  • Large file support (streaming)
  • Memory optimization
  • Parallel processing support

6.4 Error Handling

  • Detailed error messages
  • HTTP status codes (400, 404, 500, etc.)
  • Validation errors
  • Logging (Serilog recommended)

7. Test Strategy

Unit Tests

  • Separate test for each function
  • Various PDF versions
  • Edge cases (empty file, corrupt PDF, etc.)

Integration Tests

  • API endpoint tests
  • File upload/download tests
  • End-to-end scenarios

Test Files

  • Valid PDF
  • Valid PDF/A (1b, 2b, 3b)
  • Corrupt PDF
  • PDF with attachments
  • PDF with forms
  • Encrypted PDF
  • Large files (>100MB)

8. Documentation Requirements

Swagger/OpenAPI

  • Detailed description for all endpoints
  • Request/Response examples
  • Error codes description
  • Authentication (if applicable)

Code Documentation

  • XML comments (for C#)
  • Interface documentation
  • Usage examples

References


Last Updated: 03.07.2026
Prepared by: Hakan Tek
Status: Initial Draft