diff --git a/docs/api/cli.md b/docs/api/cli.md index 47f06ae..d978a3d 100644 --- a/docs/api/cli.md +++ b/docs/api/cli.md @@ -1,12 +1,36 @@ # CLI API -The command-line interface module provides functions for the tzst CLI tool. +The command-line interface module provides comprehensive functionality for the tzst CLI tool, including argument parsing, command execution, and interactive features. ```{eval-rst} .. automodule:: tzst.cli - :no-index: + :members: + :undoc-members: + :show-inheritance: ``` +## Overview + +The tzst CLI provides a powerful command-line interface for archive operations with intuitive commands and comprehensive options. The interface is designed for both interactive use and scripting, with robust error handling and user-friendly output. + +### 🔧 Core Commands + +| Command | Aliases | Description | Streaming Support | +|---------|---------|-------------|-------------------| +| `a` | `add`, `create` | Create or add to archive | N/A | +| `x` | `extract` | Extract with full paths | ✓ `--streaming` | +| `e` | `extract-flat` | Extract without directory structure | ✓ `--streaming` | +| `l` | `list` | List archive contents | ✓ `--streaming` | +| `t` | `test` | Test archive integrity | ✓ `--streaming` | + +### 🎯 Key Features + +- **Intuitive Commands**: Simple, memorable command aliases (a, x, e, l, t) +- **Streaming Support**: Memory-efficient processing for large archives +- **Interactive Conflict Resolution**: User-friendly prompts for handling file conflicts +- **Comprehensive Options**: Fine-grained control over compression, extraction, and security +- **Cross-Platform**: Consistent behavior across Windows, macOS, and Linux + ## Main Functions ### main @@ -15,12 +39,118 @@ The command-line interface module provides functions for the tzst CLI tool. .. autofunction:: tzst.cli.main ``` +The main entry point for the CLI application. Handles argument parsing, command execution, and comprehensive error reporting. + +**Key Features:** + +- Robust argument validation and error handling +- Support for all archive operations +- Consistent exit codes for scripting +- User-friendly error messages + +**Exit Codes:** + +- `0`: Success +- `1`: General error (file not found, archive corruption, etc.) +- `2`: Argument parsing error +- `130`: Interrupted by user (Ctrl+C) + ### create_parser ```{eval-rst} .. autofunction:: tzst.cli.create_parser ``` +Creates and configures the comprehensive argument parser for the CLI interface. + +**Supported Arguments:** + +- **Global**: `--version`, `--help` +- **Archive Creation**: `-l/--level`, `--no-atomic` +- **Extraction**: `-o/--output`, `--streaming`, `--filter`, `--conflict-resolution` +- **Listing**: `-v/--verbose`, `--streaming` +- **Testing**: `--streaming` + +## Command Handlers + +The CLI implements dedicated command handlers for each operation, providing specialized functionality and error handling. + +### Archive Creation Commands + +#### cmd_add + +Creates new archives from files and directories with configurable compression and atomic operations. + +**Features:** + +- Configurable compression levels (1-22) +- Atomic file operations (default) for safe creation +- Recursive directory processing +- Path validation and normalization + +**Usage Examples:** + +```bash +# Basic archive creation +tzst a backup.tzst documents/ photos/ + +# High compression with atomic disabled +tzst a backup.tzst files/ -l 15 --no-atomic +``` + +### Extraction Commands + +#### cmd_extract_full + +Extracts archives preserving complete directory structure with advanced conflict resolution. + +**Features:** + +- Preserves full directory paths +- Multiple conflict resolution strategies +- Security filters for safe extraction +- Selective file extraction +- Streaming mode for large archives + +#### cmd_extract_flat + +Extracts archives flattening all files to a single directory, useful for consolidating files. + +**Features:** + +- Flattens directory structure +- Automatic conflict resolution for filename collisions +- Preserves file content while simplifying structure +- Same security and streaming features as full extraction + +### Management Commands + +#### cmd_list + +Lists archive contents with optional detailed information and streaming support. + +**Features:** + +- Simple or verbose listing modes +- Human-readable file sizes +- Modification timestamps +- Streaming mode for memory efficiency + +#### cmd_test + +Tests archive integrity and validity with comprehensive error reporting. + +**Features:** + +- Complete archive validation +- Streaming mode support +- Detailed error reporting +- Exit codes for automated testing + +#### cmd_version + +Displays version information and system details. + ## Utility Functions ### print_banner @@ -29,14 +159,47 @@ The command-line interface module provides functions for the tzst CLI tool. .. autofunction:: tzst.cli.print_banner ``` +Displays the application banner with version and copyright information. + ### format_size ```{eval-rst} .. autofunction:: tzst.cli.format_size ``` +Formats file sizes in a human-readable format (bytes, KB, MB, GB). + ### validate_compression_level ```{eval-rst} .. autofunction:: tzst.cli.validate_compression_level ``` + +Validates compression level arguments and converts them to integers. + +## Interactive Features + +The CLI includes interactive conflict resolution for file extraction conflicts, allowing users to choose how to handle existing files during extraction operations. + +### Conflict Resolution Options + +- **Replace**: Overwrite the existing file +- **Skip**: Keep the existing file, skip extraction +- **Replace All**: Apply replace to all subsequent conflicts +- **Skip All**: Apply skip to all subsequent conflicts +- **Auto-rename All**: Automatically rename conflicting files +- **Exit**: Stop extraction process + +### Security Considerations + +The CLI implements multiple security filters for safe extraction: + +- **`data` filter** (default): Safest option, blocks potentially dangerous archive members +- **`tar` filter**: Preserves more tar features while maintaining basic security +- **`fully_trusted` filter**: No restrictions, use only with completely trusted archives + +### Performance Options + +- **Streaming Mode**: Use `--streaming` for memory-efficient processing of large archives (>100MB) +- **Compression Levels**: Choose from 1 (fastest) to 22 (maximum compression) +- **Atomic Operations**: Default behavior uses temporary files for safe archive creation diff --git a/docs/api/core.md b/docs/api/core.md index 8bc128e..911bee1 100644 --- a/docs/api/core.md +++ b/docs/api/core.md @@ -1,6 +1,6 @@ # Core API -The core module provides the main functionality for working with tzst archives. +The core module provides the main functionality for working with tzst archives, including the primary `TzstArchive` class and high-level convenience functions. ```{eval-rst} .. automodule:: tzst.core @@ -11,7 +11,7 @@ The core module provides the main functionality for working with tzst archives. ## TzstArchive Class -The main class for handling `.tzst`/`.tar.zst` archives. +The main class for handling `.tzst`/`.tar.zst` archives with comprehensive functionality for creation, extraction, and manipulation. ```{eval-rst} .. autoclass:: tzst.TzstArchive @@ -21,9 +21,32 @@ The main class for handling `.tzst`/`.tar.zst` archives. :special-members: __init__, __enter__, __exit__ ``` +### Key Features + +- **Context Manager Support**: Use with `with` statements for automatic resource management +- **Multiple Access Modes**: Read ('r'), write ('w'), and append ('a') modes +- **Streaming Support**: Memory-efficient processing for large archives +- **Security Features**: Built-in protection against path traversal attacks +- **Flexible Extraction**: Support for selective extraction and conflict resolution + +### Usage Examples + +```python +# Create a new archive +with TzstArchive("backup.tzst", "w", compression_level=6) as archive: + archive.add("important_file.txt") + archive.add("documents/", recursive=True) + +# Read an existing archive +with TzstArchive("backup.tzst", "r") as archive: + contents = archive.list(verbose=True) + is_valid = archive.test() + archive.extractall("restore/") +``` + ## Convenience Functions -High-level functions for common archive operations. +High-level functions for common archive operations without needing to instantiate the `TzstArchive` class directly. ### create_archive @@ -31,20 +54,73 @@ High-level functions for common archive operations. .. autofunction:: tzst.create_archive ``` +Creates a new tzst archive from the specified files and directories. + +**Key Features:** + +- Configurable compression levels (1-22) +- Atomic creation using temporary files +- Automatic path validation and normalization +- Support for both files and directories + ### extract_archive ```{eval-rst} .. autofunction:: tzst.extract_archive ``` +Extracts files from a tzst archive with advanced options for handling conflicts and filtering. + +**Key Features:** + +- Selective extraction with member filtering +- Multiple conflict resolution strategies +- Flatten option to extract all files to a single directory +- Streaming mode for memory efficiency +- Security filters to prevent path traversal attacks + ### list_archive ```{eval-rst} .. autofunction:: tzst.list_archive ``` +Lists the contents of a tzst archive with optional detailed information. + +**Returns:** + +- List of dictionaries containing file information +- Each entry includes name, size, modification time, and type +- Verbose mode provides additional metadata + ### test_archive ```{eval-rst} .. autofunction:: tzst.test_archive ``` + +Tests the integrity of a tzst archive to verify it can be successfully decompressed. + +**Returns:** + +- `True` if the archive is valid and can be extracted +- `False` if the archive is corrupted or cannot be processed + +## Enums and Supporting Classes + +### ConflictResolution + +Enumeration for handling file conflicts during extraction: + +- `REPLACE`: Overwrite existing files +- `SKIP`: Skip existing files +- `REPLACE_ALL`: Overwrite all existing files without prompting +- `SKIP_ALL`: Skip all existing files without prompting +- `AUTO_RENAME`: Automatically rename conflicting files +- `AUTO_RENAME_ALL`: Automatically rename all conflicting files +- `ASK`: Prompt user for each conflict (interactive mode) +- `EXIT`: Stop extraction on first conflict + +### ConflictResolutionState + +State management class for tracking conflict resolution decisions during batch operations. diff --git a/docs/api/exceptions.md b/docs/api/exceptions.md index e1b31e2..71551cf 100644 --- a/docs/api/exceptions.md +++ b/docs/api/exceptions.md @@ -1,20 +1,254 @@ -# Exceptions +# Exceptions API -Custom exception classes used by tzst. +Custom exception classes used by tzst for comprehensive error handling and debugging. ```{eval-rst} .. automodule:: tzst.exceptions - :no-index: + :members: + :undoc-members: + :show-inheritance: ``` -## Exception Hierarchy +## Overview + +The tzst library provides a comprehensive hierarchy of exceptions to help identify and handle different types of errors that may occur during archive operations. All exceptions inherit from the base `TzstError` class, making it easy to catch all tzst-related errors with a single exception handler. + +### Exception Hierarchy + +```text +TzstError (base exception) +├── TzstArchiveError (archive operation failures) +├── TzstCompressionError (compression failures) +└── TzstDecompressionError (decompression failures) +``` + +## Exception Classes + +### Base Exception + +#### TzstError + +```{eval-rst} +.. autoexception:: tzst.exceptions.TzstError + :members: + :show-inheritance: +``` + +The base exception class for all tzst operations. Catch this exception to handle any tzst-related error in your application. + +**Usage:** + +```python +from tzst import create_archive, TzstError + +try: + create_archive("backup.tzst", ["files/"]) +except TzstError as e: + print(f"tzst operation failed: {e}") +``` + +### Archive Operation Exceptions + +#### TzstArchiveError ```{eval-rst} .. autoexception:: tzst.exceptions.TzstArchiveError :members: :show-inheritance: +``` +Raised when archive operations fail, such as: + +- Archive file cannot be opened or created +- File permissions prevent archive access +- Archive structure is malformed +- Tar operations fail within the archive +- Atomic file operations fail during creation + +**Common Scenarios:** + +- Invalid archive file path +- Insufficient disk space +- File permission errors +- Corrupt archive structure + +### Compression Exceptions + +#### TzstCompressionError + +```{eval-rst} +.. autoexception:: tzst.exceptions.TzstCompressionError + :members: + :show-inheritance: +``` + +Raised when compression operations fail, including: + +- Invalid compression level is specified +- Disk space is insufficient during compression +- Input data cannot be compressed due to corruption +- Zstandard compression encounters an internal error + +**Common Scenarios:** + +- Compression level out of range (1-22) +- Insufficient disk space during compression +- Source file corruption +- Zstandard library errors + +### Decompression Exceptions + +#### TzstDecompressionError + +```{eval-rst} .. autoexception:: tzst.exceptions.TzstDecompressionError :members: :show-inheritance: ``` + +Raised when decompression operations fail, such as: + +- Archive file is corrupted or incomplete +- Archive was not created with zstandard compression +- Decompression buffer overflows or underflows +- Archive format is invalid or unsupported + +**Common Scenarios:** + +- Corrupted or truncated archive files +- Non-zstandard compressed archives +- Invalid tar structure within archive +- Archive format version mismatches + +## Error Handling Best Practices + +### Basic Error Handling + +```python +from tzst import create_archive, TzstArchiveError, TzstCompressionError + +try: + create_archive("backup.tzst", ["documents/"]) +except TzstCompressionError as e: + print(f"Compression failed: {e}") +except TzstArchiveError as e: + print(f"Archive operation failed: {e}") +``` + +### Comprehensive Error Handling + +```python +from tzst import extract_archive, TzstError + +try: + extract_archive("backup.tzst", "restore/") +except TzstError as e: + # Catch any tzst-related error + print(f"Operation failed: {e}") + # Perform cleanup or fallback operations +``` + +### Specific Exception Handling + +```python +from tzst import TzstArchive, TzstDecompressionError, TzstArchiveError + +def safe_extract(archive_path, output_dir): + try: + with TzstArchive(archive_path, "r") as archive: + # Test integrity first + if not archive.test(): + print("Archive integrity check failed") + return False + + # Extract files + archive.extractall(output_dir) + return True + + except TzstDecompressionError as e: + print(f"Archive is corrupted or invalid: {e}") + return False + except TzstArchiveError as e: + print(f"Archive operation failed: {e}") + return False + except FileNotFoundError: print(f"Archive file not found: {archive_path}") + return False + except PermissionError: + print(f"Permission denied accessing: {archive_path}") + return False + +``` + +### Logging Integration + +```python +import logging +from tzst import test_archive, TzstDecompressionError, TzstError + +logger = logging.getLogger(__name__) + +def verify_archive(archive_path): + """Verify archive integrity with comprehensive logging.""" + try: + if test_archive(archive_path): + logger.info(f"Archive {archive_path} is valid") + return True + except TzstDecompressionError as e: + logger.error(f"Archive {archive_path} is corrupted: {e}") + except TzstError as e: + logger.error(f"tzst error for {archive_path}: {e}") + except Exception as e: + logger.error(f"Unexpected error testing {archive_path}: {e}") + + return False +``` + +### Error Recovery Patterns + +```python +from tzst import create_archive, extract_archive, TzstError +from pathlib import Path +import tempfile +import shutil + +def robust_backup_and_restore(source_dir, backup_path, restore_dir): + """Robust backup with error recovery and validation.""" + temp_backup = None + + try: + # Create backup with temporary file for atomicity + with tempfile.NamedTemporaryFile(suffix='.tzst', delete=False) as temp_file: + temp_backup = Path(temp_file.name) + + # Create archive + create_archive(temp_backup, [source_dir], compression_level=6) + + # Verify archive before moving to final location + if not test_archive(temp_backup): + raise TzstArchiveError("Created archive failed integrity check") + + # Move to final location atomically + shutil.move(temp_backup, backup_path) + temp_backup = None # Successfully moved + + # Test restoration + extract_archive(backup_path, restore_dir) + + print(f"Backup and restore completed successfully") + return True + + except TzstError as e: + print(f"tzst operation failed: {e}") + # Cleanup and recovery logic + if restore_dir.exists(): + shutil.rmtree(restore_dir) + return False + except Exception as e: + print(f"Unexpected error: {e}") + return False + + finally: + # Cleanup temporary files + if temp_backup and temp_backup.exists(): + temp_backup.unlink() +``` diff --git a/docs/api/index.md b/docs/api/index.md index 1e9c515..e8e6d79 100644 --- a/docs/api/index.md +++ b/docs/api/index.md @@ -1,6 +1,6 @@ # API Reference -This section contains the complete API documentation for tzst. +This section contains the complete API documentation for tzst, providing detailed information about classes, functions, and exceptions. ```{toctree} :maxdepth: 2 @@ -12,13 +12,22 @@ exceptions ## Overview -The tzst library provides both high-level convenience functions and a comprehensive class-based API for working with `.tzst`/`.tar.zst` archives. +The tzst library provides both high-level convenience functions and a comprehensive class-based API for working with `.tzst`/`.tar.zst` archives. The library is designed with security, performance, and ease of use in mind. ### Main Components -- **{doc}`core`**: Core functionality including `TzstArchive` class and convenience functions -- **{doc}`cli`**: Command-line interface functions and utilities -- **{doc}`exceptions`**: Custom exception classes for error handling +- **{doc}`core`**: Core functionality including `TzstArchive` class and convenience functions for archive operations +- **{doc}`cli`**: Command-line interface functions and utilities for batch operations +- **{doc}`exceptions`**: Custom exception classes for comprehensive error handling and debugging + +### Architecture Overview + +The tzst library follows a layered architecture: + +1. **High-Level API**: Convenience functions for common operations +2. **Class-Based API**: `TzstArchive` class for advanced control +3. **CLI Interface**: Command-line tools for interactive and scripted use +4. **Exception System**: Comprehensive error handling for robust applications ### Quick Reference @@ -33,6 +42,8 @@ The tzst library provides both high-level convenience functions and a comprehens TzstArchive ``` +The main class for archive manipulation with context manager support and comprehensive functionality. + #### Convenience Functions ```{eval-rst} @@ -45,7 +56,26 @@ The tzst library provides both high-level convenience functions and a comprehens test_archive ``` -#### Exceptions +High-level functions that provide simple interfaces for common archive operations. + +#### CLI Functions + +```{eval-rst} +.. currentmodule:: tzst.cli + +.. autosummary:: + :nosignatures: + + main + create_parser + print_banner + format_size + validate_compression_level +``` + +Command-line interface utilities for interactive and batch operations. + +#### Exception Classes ```{eval-rst} .. currentmodule:: tzst.exceptions @@ -53,6 +83,37 @@ The tzst library provides both high-level convenience functions and a comprehens .. autosummary:: :nosignatures: + TzstError TzstArchiveError + TzstCompressionError TzstDecompressionError ``` + +Exception hierarchy for comprehensive error handling and debugging support. + +## Key Features + +### 🛡️ Security First + +- Built-in path traversal protection +- Multiple security filter options +- Safe extraction by default + +### ⚡ High Performance + +- Zstandard compression with configurable levels +- Streaming support for large archives +- Memory-efficient operations + +### 🔧 Developer Friendly + +- Clean, Pythonic API +- Comprehensive error handling +- Context manager support +- Extensive documentation and examples + +### 🌐 Cross-Platform + +- Works on Windows, macOS, and Linux +- Consistent behavior across platforms +- Native performance optimizations diff --git a/docs/examples.md b/docs/examples.md index 902294a..3396f52 100644 --- a/docs/examples.md +++ b/docs/examples.md @@ -1,523 +1,1123 @@ # Examples -This page provides practical examples of using tzst in various scenarios. +This section provides comprehensive examples of using tzst for various scenarios and use cases. + +## Table of Contents + +- [Basic Operations](#basic-operations) +- [Advanced Archive Creation](#advanced-archive-creation) +- [Flexible Extraction](#flexible-extraction) +- [Security and Filtering](#security-and-filtering) +- [Performance Optimization](#performance-optimization) +- [Error Handling](#error-handling) +- [Real-World Scenarios](#real-world-scenarios) +- [Integration Examples](#integration-examples) ## Basic Operations ### Creating Your First Archive -```python -from tzst import TzstArchive - -# Create a simple archive -with TzstArchive("my_first_archive.tzst", "w") as archive: - archive.add("important_file.txt") - archive.add("documents/", recursive=True) - -print("Archive created successfully!") -``` - -### Extracting an Archive - -```python -from tzst import TzstArchive - -# Extract everything safely -with TzstArchive("my_first_archive.tzst", "r") as archive: - archive.extract("extracted_files/", filter="data") - -print("Files extracted to extracted_files/") -``` - -## Advanced Usage - -### High-Compression Backup - ```python from tzst import create_archive -import os -# Create a highly compressed backup -def create_backup(source_dirs, backup_name): - create_archive( - archive_path=f"{backup_name}.tzst", - files=source_dirs, - compression_level=15, # High compression - ) - - # Check the compression ratio - original_size = sum( - os.path.getsize(os.path.join(dirpath, filename)) - for directory in source_dirs - if os.path.exists(directory) - for dirpath, dirnames, filenames in os.walk(directory) - for filename in filenames - ) - - compressed_size = os.path.getsize(f"{backup_name}.tzst") - ratio = (1 - compressed_size / original_size) * 100 - - print(f"Backup created: {backup_name}.tzst") - print(f"Compression ratio: {ratio:.1f}%") - print(f"Original size: {original_size:,} bytes") - print(f"Compressed size: {compressed_size:,} bytes") +# Simple archive creation +files_to_archive = ["document.pdf", "photos/", "config.json"] +create_archive("my-archive.tzst", files_to_archive) -# Usage -create_backup(["documents/", "photos/", "projects/"], "full_backup") +# With custom compression level +create_archive("high-compression.tzst", files_to_archive, compression_level=9) ``` -### Processing Large Archives with Streaming +### Command Line Equivalent + +```bash +# Create archive +tzst a my-archive.tzst document.pdf photos/ config.json + +# With high compression +tzst a high-compression.tzst document.pdf photos/ config.json --compression-level 9 +``` + +### Basic Extraction + +```python +from tzst import extract_archive + +# Extract to current directory +extract_archive("my-archive.tzst") + +# Extract to specific directory +extract_archive("my-archive.tzst", "extracted/") + +# Extract specific files only +extract_archive("my-archive.tzst", "output/", members=["document.pdf", "config.json"]) + +# Extract with conflict resolution +extract_archive("my-archive.tzst", "output/", conflict_resolution="skip") +``` + +### Listing Archive Contents + +```python +from tzst import list_archive + +# Simple listing +contents = list_archive("my-archive.tzst") +for item in contents: + print(f"{item['name']} ({item['size']} bytes)") + +# Detailed listing with timestamps and permissions +contents = list_archive("my-archive.tzst", verbose=True) +for item in contents: + print(f"{item['name']:30} {item['size']:>10} bytes {item['mtime']}") +``` + +## Command Line Usage + +### Archive Creation Commands + +```bash +# Basic archive creation +tzst a backup.tzst documents/ photos/ + +# Create with specific compression level +tzst a backup.tzst documents/ photos/ -l 6 + +# Create from multiple sources +tzst a complete-backup.tzst /home/user/documents /home/user/photos /etc/config + +# Create with verbose output +tzst a backup.tzst documents/ photos/ -v +``` + +### Extraction Commands + +```bash +# Extract to current directory +tzst x backup.tzst + +# Extract to specific directory +tzst x backup.tzst --output /restore/ + +# Extract specific files only +tzst x backup.tzst documents/report.pdf photos/vacation.jpg + +# Extract with conflict resolution +tzst x backup.tzst --conflict-resolution skip + +# Extract flattening directory structure +tzst e backup.tzst --output flat-restore/ +``` + +### Archive Inspection Commands + +```bash +# List archive contents +tzst l backup.tzst + +# List with detailed information +tzst l backup.tzst --verbose + +# Test archive integrity +tzst t backup.tzst + +# Stream large archives efficiently +tzst l huge-archive.tzst --streaming +``` + +## Advanced Archive Creation + +### Working with the TzstArchive Class ```python from tzst import TzstArchive +from pathlib import Path -def process_large_archive(archive_path, output_dir): - """Process a large archive efficiently using streaming mode.""" +# Create archive with fine-grained control +with TzstArchive("project-backup.tzst", "w", compression_level=6) as archive: + # Add individual files + archive.add("README.md") + archive.add("LICENSE") - # Use streaming to handle large archives - with TzstArchive(archive_path, "r", streaming=True) as archive: - # First, list contents to understand what we're dealing with - print("Analyzing archive contents...") - contents = archive.list(verbose=True) - - total_files = sum(1 for item in contents if item['is_file']) - total_size = sum(item['size'] for item in contents if item['is_file']) - - print(f"Archive contains {total_files} files ({total_size:,} bytes)") - - # Extract only specific file types - text_files = [item['name'] for item in contents - if item['name'].endswith(('.txt', '.md', '.py'))] - - if text_files: - print(f"Extracting {len(text_files)} text files...") - archive.extract(output_dir, members=text_files, filter="data") - - print("Processing complete!") - -# Usage -process_large_archive("large_dataset.tzst", "extracted_text_files/") + # Add directories recursively + archive.add("src/", recursive=True) + archive.add("tests/", recursive=True) + + # Add with custom archive names + archive.add("config/production.yaml", arcname="config.yaml") + archive.add("/tmp/build-info.json", arcname="build-info.json") ``` -### Batch Archive Operations +### Conditional File Addition ```python -from tzst import test_archive, list_archive import os -from pathlib import Path - -def verify_archive_collection(archive_dir): - """Verify integrity of all archives in a directory.""" - - archive_dir = Path(archive_dir) - archives = list(archive_dir.glob("*.tzst")) - - print(f"Found {len(archives)} archives to verify...") - - results = [] - for archive_path in archives: - print(f"Testing {archive_path.name}...") - - try: - # Test integrity - is_valid = test_archive(str(archive_path)) - - if is_valid: - # Get archive info - contents = list_archive(str(archive_path), verbose=True) - file_count = sum(1 for item in contents if item['is_file']) - total_size = sum(item['size'] for item in contents if item['is_file']) - - results.append({ - 'name': archive_path.name, - 'status': 'Valid', - 'file_count': file_count, - 'total_size': total_size - }) - else: - results.append({ - 'name': archive_path.name, - 'status': 'Corrupted', - 'file_count': 0, - 'total_size': 0 - }) - - except Exception as e: - results.append({ - 'name': archive_path.name, - 'status': f'Error: {e}', - 'file_count': 0, - 'total_size': 0 - }) - - # Print summary - print("\n" + "="*60) - print("VERIFICATION SUMMARY") - print("="*60) - - for result in results: - status = result['status'] - if status == 'Valid': - print(f"✓ {result['name']}: {result['file_count']} files, " - f"{result['total_size']:,} bytes") - else: - print(f"✗ {result['name']}: {status}") - - valid_count = sum(1 for r in results if r['status'] == 'Valid') - print(f"\nSummary: {valid_count}/{len(results)} archives are valid") - -# Usage -verify_archive_collection("backup_archives/") -``` - -## Security Examples - -### Safe Archive Extraction - -```python from tzst import TzstArchive -from tzst.exceptions import TzstArchiveError - -def safe_extract(archive_path, output_dir, max_size_mb=100): - """Safely extract an archive with size limits and security filters.""" - - try: - with TzstArchive(archive_path, "r") as archive: - # First, analyze the archive - contents = archive.list(verbose=True) - - # Check total uncompressed size - total_size = sum(item['size'] for item in contents if item['is_file']) - max_size_bytes = max_size_mb * 1024 * 1024 - - if total_size > max_size_bytes: - print(f"Warning: Archive is {total_size:,} bytes when uncompressed") - print(f"This exceeds the limit of {max_size_bytes:,} bytes") - response = input("Continue anyway? (y/N): ") - if response.lower() != 'y': - return False - - # Check for suspicious files - suspicious_files = [] - for item in contents: - name = item['name'] - # Check for directory traversal attempts - if '..' in name or name.startswith('/'): - suspicious_files.append(name) - # Check for executable files - if name.endswith(('.exe', '.bat', '.sh', '.com')): - suspicious_files.append(name) - - if suspicious_files: - print(f"Warning: Found {len(suspicious_files)} suspicious files:") - for file in suspicious_files[:5]: # Show first 5 - print(f" - {file}") - if len(suspicious_files) > 5: - print(f" ... and {len(suspicious_files) - 5} more") - - response = input("Continue extraction? (y/N): ") - if response.lower() != 'y': - return False - - # Extract with the safest filter - print(f"Extracting {len(contents)} items to {output_dir}") - archive.extract(output_dir, filter="data") - print("Extraction completed safely!") - return True - - except TzstArchiveError as e: - print(f"Archive error: {e}") - return False - except Exception as e: - print(f"Unexpected error: {e}") - return False - -# Usage -safe_extract("untrusted_archive.tzst", "safe_output/", max_size_mb=50) -``` - -### Archive Validation Pipeline - -```python -from tzst import TzstArchive, test_archive -import hashlib -import json from pathlib import Path -def create_archive_manifest(archive_path): - """Create a manifest of archive contents for validation.""" +def backup_project(project_path, output_archive): + """Create a project backup excluding certain files.""" + project_path = Path(project_path) - manifest = { - 'archive_path': str(archive_path), - 'files': [], - 'created_at': str(Path(archive_path).stat().st_mtime) + # Define exclusion patterns + exclude_patterns = { + "*.pyc", "*.pyo", "__pycache__", + ".git", ".svn", "node_modules", + "*.tmp", "*.log", ".DS_Store" } - with TzstArchive(archive_path, "r") as archive: - contents = archive.list(verbose=True) - - for item in contents: - if item['is_file']: - manifest['files'].append({ - 'name': item['name'], - 'size': item['size'], - 'mtime': item['mtime'] - }) - - # Save manifest - manifest_path = Path(archive_path).with_suffix('.manifest.json') - with open(manifest_path, 'w') as f: - json.dump(manifest, f, indent=2) - - print(f"Manifest created: {manifest_path}") - return manifest - -def validate_archive_with_manifest(archive_path): - """Validate an archive against its manifest.""" - - manifest_path = Path(archive_path).with_suffix('.manifest.json') - - if not manifest_path.exists(): - print("No manifest found, creating new one...") - create_archive_manifest(archive_path) - return True - - # Load manifest - with open(manifest_path, 'r') as f: - manifest = json.load(f) - - print("Validating archive integrity...") - if not test_archive(archive_path): - print("❌ Archive integrity check failed!") - return False - - print("Validating against manifest...") - with TzstArchive(archive_path, "r") as archive: - contents = archive.list(verbose=True) - current_files = {item['name']: item for item in contents if item['is_file']} - - manifest_files = {item['name']: item for item in manifest['files']} - - # Check for missing files - missing = set(manifest_files.keys()) - set(current_files.keys()) - if missing: - print(f"❌ Missing files: {', '.join(missing)}") - return False - - # Check for extra files - extra = set(current_files.keys()) - set(manifest_files.keys()) - if extra: - print(f"⚠️ Extra files: {', '.join(extra)}") - - # Check file sizes - size_mismatches = [] - for name, manifest_file in manifest_files.items(): - if name in current_files: - if current_files[name]['size'] != manifest_file['size']: - size_mismatches.append(name) - - if size_mismatches: - print(f"❌ Size mismatches: {', '.join(size_mismatches)}") - return False - - print("✅ Archive validation passed!") - return True + with TzstArchive(output_archive, "w", compression_level=5) as archive: + for item in project_path.rglob("*"): + # Skip excluded patterns + if any(item.match(pattern) for pattern in exclude_patterns): + continue + + # Skip if it's a directory (will be created automatically) + if item.is_dir(): + continue + + # Add file with relative path + rel_path = item.relative_to(project_path) + archive.add(str(item), arcname=str(rel_path)) + print(f"Added: {rel_path}") # Usage -archive_path = "important_backup.tzst" -if validate_archive_with_manifest(archive_path): - print("Archive is valid and matches manifest") -else: - print("Archive validation failed!") +backup_project("/home/user/myproject", "project-clean.tzst") ``` -## Performance Examples - -### Compression Level Comparison +### Atomic Archive Creation ```python from tzst import create_archive -import time + +# Safe atomic creation (default behavior) +# Creates in temporary file first, then moves to final location +create_archive("important-data.tzst", ["critical/"], use_temp_file=True) + +# Direct creation (faster but not atomic) +create_archive("temp-data.tzst", ["temp/"], use_temp_file=False) +``` + +## Flexible Extraction + +### Extracting with Different Structures + +```python +from tzst import extract_archive, TzstArchive + +# Standard extraction (preserves directory structure) +extract_archive("archive.tzst", "output/") + +# Flatten all files to single directory +extract_archive("archive.tzst", "flat-output/", flatten=True) + +# Extract with streaming for large archives +extract_archive("huge-archive.tzst", "output/", streaming=True) +``` + +### Selective Extraction + +```python +from tzst import TzstArchive + +def extract_by_extension(archive_path, output_dir, extensions): + """Extract only files with specific extensions.""" + with TzstArchive(archive_path, "r") as archive: + members = archive.getmembers() + + # Filter members by extension + filtered_members = [ + member.name for member in members + if any(member.name.endswith(ext) for ext in extensions) + ] + + if filtered_members: + archive.extractall(output_dir, members=filtered_members) + print(f"Extracted {len(filtered_members)} files") + else: + print("No matching files found") + +# Extract only images +extract_by_extension("photos.tzst", "images/", [".jpg", ".png", ".gif"]) + +# Extract only documents +extract_by_extension("backup.tzst", "docs/", [".pdf", ".docx", ".txt"]) +``` + +### Custom Extraction Logic + +```python +from tzst import TzstArchive import os + +def extract_large_files_only(archive_path, output_dir, min_size_mb=10): + """Extract only files larger than specified size.""" + min_size_bytes = min_size_mb * 1024 * 1024 + + with TzstArchive(archive_path, "r") as archive: + large_files = [] + + for member in archive.getmembers(): + if member.isfile() and member.size > min_size_bytes: + large_files.append(member.name) + size_mb = member.size / (1024 * 1024) + print(f"Will extract: {member.name} ({size_mb:.1f} MB)") + + if large_files: + os.makedirs(output_dir, exist_ok=True) + for filename in large_files: + archive.extract(filename, output_dir) + print(f"Extracted {len(large_files)} large files") + +extract_large_files_only("mixed-content.tzst", "large-files/", min_size_mb=5) +``` + +## Security and Filtering + +### Safe Extraction Practices + +```python +from tzst import extract_archive + +# Always use secure filters (default behavior) +extract_archive("untrusted.tzst", "safe-output/", filter="data") + +# For trusted archives with special tar features +extract_archive("trusted.tzst", "output/", filter="tar") + +# Only for completely trusted archives +extract_archive("internal.tzst", "output/", filter="fully_trusted") +``` + +### Custom Security Filter + +```python +import tarfile +from tzst import TzstArchive + +def secure_data_filter(member, path): + """Custom filter that only allows regular files and directories.""" + # Only allow regular files and directories + if not (member.isfile() or member.isdir()): + return None + + # Prevent path traversal + if os.path.isabs(member.name) or ".." in member.name: + return None + + # Limit file size (100MB max) + if member.isfile() and member.size > 100 * 1024 * 1024: + return None + + return member + +# Use custom filter +with TzstArchive("archive.tzst", "r") as archive: + archive.extractall("secure-output/", filter=secure_data_filter) +``` + +### Handling File Conflicts + +```python +from tzst import extract_archive, ConflictResolution + +# Skip existing files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.SKIP_ALL) + +# Replace all existing files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.REPLACE_ALL) + +# Auto-rename conflicting files (adds suffix like "_1", "_2", etc.) +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.AUTO_RENAME_ALL) + +# Interactive resolution (command line only) +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.ASK) +``` + +### Custom Conflict Resolution + +```python +from tzst import extract_archive, ConflictResolution from pathlib import Path -def compression_benchmark(files, output_prefix="test"): - """Compare different compression levels for the same files.""" +def custom_conflict_handler(target_path: Path) -> ConflictResolution: + """Custom logic for handling file conflicts.""" + # Check file age + if target_path.exists(): + file_age_days = (time.time() - target_path.stat().st_mtime) / (24 * 3600) + + if file_age_days > 30: + print(f"Replacing old file: {target_path}") + return ConflictResolution.REPLACE + else: + print(f"Keeping newer file: {target_path}") + return ConflictResolution.SKIP + + return ConflictResolution.REPLACE + +# Use custom callback +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.ASK, + interactive_callback=custom_conflict_handler) +``` + +## Performance Optimization + +### Streaming for Large Archives + +```python +from tzst import TzstArchive, list_archive, test_archive + +# Memory-efficient operations for large archives +large_archive = "backup-500gb.tzst" + +# Test integrity with streaming +is_valid = test_archive(large_archive, streaming=True) + +# List contents with streaming +contents = list_archive(large_archive, streaming=True, verbose=True) + +# Extract with streaming +with TzstArchive(large_archive, "r", streaming=True) as archive: + archive.extractall("restore/") +``` + +### Compression Level Optimization + +```python +import time +from tzst import create_archive + +def benchmark_compression_levels(files, output_prefix="test"): + """Compare different compression levels.""" + levels_to_test = [1, 3, 6, 9, 15, 22] results = [] - levels = [1, 3, 6, 9, 15, 22] # Representative levels - - for level in levels: - archive_path = f"{output_prefix}_level_{level}.tzst" + for level in levels_to_test: + output_file = f"{output_prefix}_level_{level}.tzst" - print(f"Testing compression level {level}...") + # Measure compression time start_time = time.time() + create_archive(output_file, files, compression_level=level) + compress_time = time.time() - start_time - create_archive( - archive_path=archive_path, - files=files, - compression_level=level - ) - - compression_time = time.time() - start_time - archive_size = os.path.getsize(archive_path) + # Get file size + file_size = Path(output_file).stat().st_size results.append({ 'level': level, - 'time': compression_time, - 'size': archive_size, - 'path': archive_path + 'time': compress_time, + 'size': file_size, + 'size_mb': file_size / (1024 * 1024) }) - print(f" Time: {compression_time:.2f}s, Size: {archive_size:,} bytes") - - # Print comparison table - print("\n" + "="*70) - print("COMPRESSION LEVEL COMPARISON") - print("="*70) - print(f"{'Level':<6} {'Time (s)':<10} {'Size (MB)':<12} {'Ratio':<8} {'Speed'}") - print("-" * 70) - - baseline_size = results[0]['size'] # Level 1 as baseline - baseline_time = results[0]['time'] - - for result in results: - size_mb = result['size'] / (1024 * 1024) - ratio = result['size'] / baseline_size - speed_factor = baseline_time / result['time'] - - print(f"{result['level']:<6} {result['time']:<10.2f} {size_mb:<12.1f} " - f"{ratio:<8.2f} {speed_factor:<.2f}x") - - # Clean up test files - for result in results: - os.remove(result['path']) + print(f"Level {level}: {compress_time:.2f}s, {file_size/1024/1024:.1f} MB") return results -# Usage -benchmark_files = ["large_directory/", "data_files/"] -results = compression_benchmark(benchmark_files, "benchmark") +# Test different compression levels +results = benchmark_compression_levels(["large-directory/"]) ``` -## Error Handling Examples - -### Robust Archive Processing +### Parallel Processing ```python -from tzst import TzstArchive, extract_archive -from tzst.exceptions import TzstArchiveError, TzstDecompressionError +import concurrent.futures +from tzst import create_archive +from pathlib import Path + +def create_archive_batch(file_groups, output_dir="archives/", compression_level=6): + """Create multiple archives in parallel.""" + Path(output_dir).mkdir(exist_ok=True) + + def create_single_archive(args): + group_name, files = args + output_path = Path(output_dir) / f"{group_name}.tzst" + create_archive(output_path, files, compression_level=compression_level) + return f"Created {output_path}" + + # Use ThreadPoolExecutor for I/O-bound operations + with concurrent.futures.ThreadPoolExecutor(max_workers=4) as executor: + future_to_group = { + executor.submit(create_single_archive, item): item[0] + for item in file_groups.items() + } + + for future in concurrent.futures.as_completed(future_to_group): + group_name = future_to_group[future] + try: + result = future.result() + print(result) + except Exception as e: + print(f"Archive {group_name} failed: {e}") + +# Example usage +file_groups = { + "documents": ["docs/", "papers/"], + "projects": ["src/", "tests/"], + "media": ["photos/", "videos/"] +} + +create_archive_batch(file_groups) +``` + +## Error Handling + +### Comprehensive Error Handling + +```python +from tzst import TzstArchive, TzstArchiveError, TzstDecompressionError import logging -# Set up logging -logging.basicConfig(level=logging.INFO) -logger = logging.getLogger(__name__) +def safe_archive_operation(operation, *args, **kwargs): + """Wrapper for safe archive operations with logging.""" + try: + return operation(*args, **kwargs) + except TzstDecompressionError as e: + logging.error(f"Decompression error: {e}") + print("The archive appears to be corrupted or not a valid tzst file.") + return None + except TzstArchiveError as e: + logging.error(f"Archive error: {e}") + print(f"Archive operation failed: {e}") + return None + except PermissionError as e: + logging.error(f"Permission error: {e}") + print("Permission denied. Check file/directory permissions.") + return None + except FileNotFoundError as e: + logging.error(f"File not found: {e}") + print(f"File or directory not found: {e}") + return None + except Exception as e: + logging.error(f"Unexpected error: {e}") + print(f"An unexpected error occurred: {e}") + return None -def robust_archive_processor(archive_paths, output_base_dir): - """Process multiple archives with comprehensive error handling.""" +# Example usage +def create_backup_safely(files, output_archive): + def create_operation(): + from tzst import create_archive + return create_archive(output_archive, files) - results = { - 'success': [], - 'failed': [], - 'skipped': [] - } + result = safe_archive_operation(create_operation) + if result is not None: + print(f"Backup created successfully: {output_archive}") + else: + print("Backup creation failed!") +``` + +### Validation and Recovery + +```python +from tzst import test_archive, list_archive, TzstArchive +from pathlib import Path + +def validate_and_repair_archive(archive_path): + """Validate archive and attempt basic recovery.""" + archive_path = Path(archive_path) - for archive_path in archive_paths: + print(f"Validating {archive_path}...") + + # Test basic integrity + try: + if test_archive(archive_path): + print("✅ Archive integrity test passed") + return True + except Exception as e: + print(f"❌ Integrity test failed: {e}") + + # Try to list contents + try: + contents = list_archive(archive_path) + print(f"📁 Archive contains {len(contents)} items") + + # Try streaming mode if regular mode fails + contents_streaming = list_archive(archive_path, streaming=True) + if len(contents_streaming) != len(contents): + print("⚠️ Different results between modes - possible corruption") + + except Exception as e: + print(f"❌ Cannot list contents: {e}") + return False + + # Try partial extraction + try: + backup_dir = archive_path.parent / f"{archive_path.stem}_recovery" + backup_dir.mkdir(exist_ok=True) + + with TzstArchive(archive_path, "r") as archive: + extracted_count = 0 + for member in archive.getmembers(): + try: + if member.isfile(): + archive.extract(member.name, backup_dir) + extracted_count += 1 + except Exception as e: + print(f"⚠️ Failed to extract {member.name}: {e}") + + print(f"✅ Recovered {extracted_count} files to {backup_dir}") + return True + + except Exception as e: + print(f"❌ Recovery failed: {e}") + return False + +# Example usage +validate_and_repair_archive("potentially-corrupted.tzst") +``` + +## Real-World Scenarios + +### Automated Backup System + +```python +#!/usr/bin/env python3 +""" +Daily backup script with rotation and validation. +""" + +import os +import sys +from datetime import datetime, timedelta +from pathlib import Path +from tzst import create_archive, test_archive + +class BackupManager: + def __init__(self, source_dirs, backup_dir, retention_days=30): + self.source_dirs = [Path(d) for d in source_dirs] + self.backup_dir = Path(backup_dir) + self.retention_days = retention_days + self.backup_dir.mkdir(parents=True, exist_ok=True) + + def create_backup(self): + """Create a new backup with timestamp.""" + timestamp = datetime.now().strftime("%Y%m%d_%H%M%S") + backup_name = f"backup_{timestamp}.tzst" + backup_path = self.backup_dir / backup_name + + print(f"Creating backup: {backup_name}") + + # Collect all existing files + files_to_backup = [] + for source_dir in self.source_dirs: + if source_dir.exists(): + files_to_backup.append(str(source_dir)) + else: + print(f"Warning: Source directory not found: {source_dir}") + + if not files_to_backup: + print("No files to backup!") + return None + try: - logger.info(f"Processing {archive_path}") + # Create backup with high compression for storage efficiency + create_archive(backup_path, files_to_backup, compression_level=9) - # Create output directory for this archive - archive_name = Path(archive_path).stem - output_dir = Path(output_base_dir) / archive_name - output_dir.mkdir(parents=True, exist_ok=True) - - # First, test the archive - logger.info(f"Testing integrity of {archive_path}") - with TzstArchive(archive_path, "r") as archive: - if not archive.test(): - raise TzstArchiveError(f"Archive {archive_path} failed integrity test") + # Validate the backup + if test_archive(backup_path): + file_size = backup_path.stat().st_size / (1024 * 1024) + print(f"✅ Backup created and validated: {file_size:.1f} MB") + return backup_path + else: + print("❌ Backup validation failed!") + backup_path.unlink() # Remove invalid backup + return None - # Get archive info - contents = archive.list(verbose=True) - file_count = sum(1 for item in contents if item['is_file']) - total_size = sum(item['size'] for item in contents if item['is_file']) - - logger.info(f"Archive contains {file_count} files ({total_size:,} bytes)") - - # Extract with error handling - logger.info(f"Extracting to {output_dir}") - archive.extract(str(output_dir), filter="data") - - results['success'].append({ - 'path': archive_path, - 'file_count': file_count, - 'total_size': total_size, - 'output_dir': str(output_dir) - }) - - logger.info(f"Successfully processed {archive_path}") - - except TzstArchiveError as e: - logger.error(f"Archive error processing {archive_path}: {e}") - results['failed'].append({ - 'path': archive_path, - 'error': str(e), - 'error_type': 'TzstArchiveError' - }) - - except TzstDecompressionError as e: - logger.error(f"Decompression error processing {archive_path}: {e}") - results['failed'].append({ - 'path': archive_path, - 'error': str(e), - 'error_type': 'TzstDecompressionError' - }) - - except FileNotFoundError: - logger.warning(f"Archive not found: {archive_path}") - results['skipped'].append({ - 'path': archive_path, - 'reason': 'File not found' - }) - - except PermissionError as e: - logger.error(f"Permission error processing {archive_path}: {e}") - results['failed'].append({ - 'path': archive_path, - 'error': str(e), - 'error_type': 'PermissionError' - }) - except Exception as e: - logger.error(f"Unexpected error processing {archive_path}: {e}") - results['failed'].append({ - 'path': archive_path, - 'error': str(e), - 'error_type': 'UnexpectedError' - }) + print(f"❌ Backup failed: {e}") + return None - # Print summary - print("\n" + "="*60) - print("PROCESSING SUMMARY") - print("="*60) - print(f"✅ Successfully processed: {len(results['success'])}") - print(f"❌ Failed: {len(results['failed'])}") - print(f"⏭️ Skipped: {len(results['skipped'])}") + def cleanup_old_backups(self): + """Remove backups older than retention period.""" + cutoff_date = datetime.now() - timedelta(days=self.retention_days) + + removed_count = 0 + for backup_file in self.backup_dir.glob("backup_*.tzst"): + # Extract timestamp from filename + try: + timestamp_str = backup_file.stem.split("_", 1)[1] + file_date = datetime.strptime(timestamp_str, "%Y%m%d_%H%M%S") + + if file_date < cutoff_date: + backup_file.unlink() + removed_count += 1 + print(f"Removed old backup: {backup_file.name}") + + except (ValueError, IndexError): + print(f"Warning: Could not parse backup date: {backup_file.name}") + + print(f"Cleaned up {removed_count} old backups") - if results['failed']: - print("\nFailures:") - for failure in results['failed']: - print(f" - {failure['path']}: {failure['error_type']}") + def run_backup(self): + """Run complete backup process.""" + print("Starting backup process...") + + backup_path = self.create_backup() + if backup_path: + self.cleanup_old_backups() + print("Backup process completed successfully!") + return True + else: + print("Backup process failed!") + return False + +# Configuration +if __name__ == "__main__": + # Customize these paths for your setup + BACKUP_SOURCES = [ + "~/Documents", + "~/Projects", + "~/Pictures", + "/etc", # System configs (Linux/macOS) + ] - return results + BACKUP_DESTINATION = "~/Backups" + RETENTION_DAYS = 30 + + # Expand user paths + sources = [os.path.expanduser(path) for path in BACKUP_SOURCES] + destination = os.path.expanduser(BACKUP_DESTINATION) + + # Run backup + backup_manager = BackupManager(sources, destination, RETENTION_DAYS) + success = backup_manager.run_backup() + + sys.exit(0 if success else 1) +``` + +### Log File Archiver + +```python +#!/usr/bin/env python3 +""" +Archive and compress log files by date. +""" + +import re +from datetime import datetime, timedelta +from pathlib import Path +from tzst import create_archive + +def archive_logs_by_date(log_dir, archive_dir, days_old=7): + """Archive log files older than specified days.""" + log_dir = Path(log_dir) + archive_dir = Path(archive_dir) + archive_dir.mkdir(parents=True, exist_ok=True) + + cutoff_date = datetime.now() - timedelta(days=days_old) + + # Group log files by date + log_groups = {} + log_pattern = re.compile(r"(\d{4}-\d{2}-\d{2})") + + for log_file in log_dir.glob("*.log"): + # Try to extract date from filename or modification time + date_match = log_pattern.search(log_file.name) + if date_match: + file_date_str = date_match.group(1) + try: + file_date = datetime.strptime(file_date_str, "%Y-%m-%d") + except ValueError: + # Fall back to modification time + file_date = datetime.fromtimestamp(log_file.stat().st_mtime) + else: + # Use modification time + file_date = datetime.fromtimestamp(log_file.stat().st_mtime) + + # Skip recent files + if file_date >= cutoff_date: + continue + + # Group by date + date_key = file_date.strftime("%Y-%m-%d") + if date_key not in log_groups: + log_groups[date_key] = [] + log_groups[date_key].append(log_file) + + # Create archives for each date group + archived_files = [] + for date_key, files in log_groups.items(): + archive_name = f"logs_{date_key}.tzst" + archive_path = archive_dir / archive_name + + # Skip if archive already exists + if archive_path.exists(): + print(f"Archive already exists: {archive_name}") + continue + + print(f"Archiving {len(files)} log files for {date_key}") + + try: + # Create archive with maximum compression (logs compress well) + create_archive(archive_path, [str(f) for f in files], compression_level=22) + + # Verify archive + from tzst import test_archive + if test_archive(archive_path): + # Remove original files after successful archiving + for log_file in files: + log_file.unlink() + archived_files.append(log_file) + + file_size = archive_path.stat().st_size / 1024 + print(f"✅ Created {archive_name} ({file_size:.1f} KB)") + else: + print(f"❌ Archive validation failed for {archive_name}") + archive_path.unlink() + + except Exception as e: + print(f"❌ Failed to archive logs for {date_key}: {e}") + + print(f"Archived {len(archived_files)} log files") # Usage -archive_list = [ - "backup1.tzst", - "backup2.tzst", - "backup3.tzst", - "missing_file.tzst" # This will be skipped -] - -results = robust_archive_processor(archive_list, "extracted_archives/") +if __name__ == "__main__": + archive_logs_by_date("/var/log", "/var/archives", days_old=7) ``` + +### Data Migration Tool + +```python +#!/usr/bin/env python3 +""" +Migrate data between systems using tzst archives. +""" + +import hashlib +from pathlib import Path +from tzst import create_archive, extract_archive, test_archive + +class DataMigrator: + def __init__(self, source_dir, staging_dir): + self.source_dir = Path(source_dir) + self.staging_dir = Path(staging_dir) + self.staging_dir.mkdir(parents=True, exist_ok=True) + + def calculate_checksum(self, file_path): + """Calculate SHA256 checksum of a file.""" + sha256_hash = hashlib.sha256() + with open(file_path, "rb") as f: + for chunk in iter(lambda: f.read(4096), b""): + sha256_hash.update(chunk) + return sha256_hash.hexdigest() + + def create_migration_package(self, package_name): + """Create a migration package with checksums.""" + package_path = self.staging_dir / f"{package_name}.tzst" + checksum_file = self.staging_dir / f"{package_name}.sha256" + + print(f"Creating migration package: {package_name}") + + # Create the archive + create_archive( + package_path, + [str(self.source_dir)], + compression_level=6 # Balanced for network transfer + ) + + # Verify archive + if not test_archive(package_path): + raise RuntimeError("Archive validation failed") + + # Calculate and save checksum + checksum = self.calculate_checksum(package_path) + with open(checksum_file, "w") as f: + f.write(f"{checksum} {package_path.name}\n") + + package_size = package_path.stat().st_size / (1024 * 1024) + print(f"✅ Package created: {package_size:.1f} MB") + print(f"📋 Checksum: {checksum}") + + return package_path, checksum_file + + def verify_and_extract_package(self, package_path, checksum_path, destination): + """Verify package integrity and extract.""" + package_path = Path(package_path) + checksum_path = Path(checksum_path) + destination = Path(destination) + + print(f"Verifying package: {package_path.name}") + + # Verify checksum + expected_checksum = checksum_path.read_text().strip().split()[0] + actual_checksum = self.calculate_checksum(package_path) + + if expected_checksum != actual_checksum: + raise RuntimeError(f"Checksum mismatch! Expected: {expected_checksum}, Got: {actual_checksum}") + + print("✅ Checksum verification passed") + + # Test archive integrity + if not test_archive(package_path): + raise RuntimeError("Archive integrity check failed") + + print("✅ Archive integrity verified") + + # Extract with conflict resolution + destination.mkdir(parents=True, exist_ok=True) + extract_archive( + package_path, + destination, + conflict_resolution="replace_all" # Overwrite for migration + ) + + print(f"✅ Package extracted to: {destination}") + +# Example usage +if __name__ == "__main__": + # Create migration package + migrator = DataMigrator("/home/user/important-data", "/tmp/migration") + package_path, checksum_path = migrator.create_migration_package("data-migration-v1") + + # Simulate transfer and extraction on target system + migrator.verify_and_extract_package( + package_path, + checksum_path, + "/home/user/restored-data" + ) +``` + +## Integration Examples + +### Django Management Command + +```python +# management/commands/backup_media.py +from django.core.management.base import BaseCommand +from django.conf import settings +from tzst import create_archive +from datetime import datetime +import os + +class Command(BaseCommand): + help = 'Create a backup of media files' + + def add_arguments(self, parser): + parser.add_argument( + '--output-dir', + default='/backups', + help='Output directory for backup files' + ) + parser.add_argument( + '--compression-level', + type=int, + default=6, + help='Compression level (1-22)' + ) + + def handle(self, *args, **options): + media_root = settings.MEDIA_ROOT + output_dir = options['output_dir'] + compression_level = options['compression_level'] + + if not os.path.exists(media_root): + self.stdout.write( + self.style.ERROR(f'Media directory not found: {media_root}') + ) + return + + # Create backup filename with timestamp + timestamp = datetime.now().strftime('%Y%m%d_%H%M%S') + backup_filename = f'media_backup_{timestamp}.tzst' + backup_path = os.path.join(output_dir, backup_filename) + + # Ensure output directory exists + os.makedirs(output_dir, exist_ok=True) + + try: + self.stdout.write(f'Creating media backup: {backup_filename}') + create_archive(backup_path, [media_root], compression_level=compression_level) + + # Verify backup + from tzst import test_archive + if test_archive(backup_path): + file_size = os.path.getsize(backup_path) / (1024 * 1024) + self.stdout.write( + self.style.SUCCESS( + f'Backup created successfully: {backup_filename} ({file_size:.1f} MB)' + ) + ) + else: + self.stdout.write( + self.style.ERROR('Backup validation failed!') + ) + + except Exception as e: + self.stdout.write( + self.style.ERROR(f'Backup failed: {e}') + ) +``` + +### Flask Application Integration + +```python +from flask import Flask, request, send_file, jsonify +from tzst import create_archive, extract_archive +import tempfile +import os +from pathlib import Path + +app = Flask(__name__) + +@app.route('/api/backup', methods=['POST']) +def create_backup(): + """API endpoint to create backups.""" + try: + data = request.get_json() + paths = data.get('paths', []) + compression_level = data.get('compression_level', 6) + + if not paths: + return jsonify({'error': 'No paths specified'}), 400 + + # Create temporary archive + with tempfile.NamedTemporaryFile(suffix='.tzst', delete=False) as tmp: + temp_path = tmp.name + + create_archive(temp_path, paths, compression_level=compression_level) + + # Return archive file + return send_file( + temp_path, + as_attachment=True, + download_name='backup.tzst', + mimetype='application/octet-stream' + ) + + except Exception as e: + return jsonify({'error': str(e)}), 500 + finally: + # Clean up temporary file + if 'temp_path' in locals() and os.path.exists(temp_path): + os.unlink(temp_path) + +@app.route('/api/extract', methods=['POST']) +def extract_files(): + """API endpoint to extract archives.""" + try: + if 'file' not in request.files: + return jsonify({'error': 'No file uploaded'}), 400 + + file = request.files['file'] + if file.filename == '': + return jsonify({'error': 'No file selected'}), 400 + + # Save uploaded file temporarily + with tempfile.NamedTemporaryFile(suffix='.tzst', delete=False) as tmp: + temp_archive = tmp.name + file.save(temp_archive) + + # Create extraction directory + extract_dir = tempfile.mkdtemp() + + # Extract archive + extract_archive(temp_archive, extract_dir) + + # List extracted files + extracted_files = [] + for root, dirs, files in os.walk(extract_dir): + for file in files: + rel_path = os.path.relpath(os.path.join(root, file), extract_dir) + extracted_files.append(rel_path) + + return jsonify({ + 'success': True, + 'extracted_files': extracted_files, + 'extract_path': extract_dir + }) + + except Exception as e: + return jsonify({'error': str(e)}), 500 + finally: + # Clean up temporary archive + if 'temp_archive' in locals() and os.path.exists(temp_archive): + os.unlink(temp_archive) + +if __name__ == '__main__': + app.run(debug=True) +``` + +### Jupyter Notebook Integration + +```python +# Cell 1: Setup +import pandas as pd +from tzst import create_archive, extract_archive, list_archive +from pathlib import Path +import matplotlib.pyplot as plt + +# Cell 2: Create dataset archive +def archive_datasets(data_dir="./data", archive_name="datasets.tzst"): + """Archive all dataset files for sharing.""" + data_path = Path(data_dir) + + if not data_path.exists(): + print(f"Creating sample data directory: {data_dir}") + data_path.mkdir(exist_ok=True) + + # Create sample datasets + sample_data = pd.DataFrame({ + 'A': range(100), + 'B': range(100, 200), + 'C': range(200, 300) + }) + + sample_data.to_csv(data_path / "sample.csv", index=False) + sample_data.to_parquet(data_path / "sample.parquet") + + # Create archive + create_archive(archive_name, [str(data_path)], compression_level=9) + + # Show archive contents + contents = list_archive(archive_name, verbose=True) + df = pd.DataFrame(contents) + + print(f"Created archive: {archive_name}") + return df + +# Execute +archive_contents = archive_datasets() +display(archive_contents) + +# Cell 3: Analyze archive +def analyze_archive(archive_path="datasets.tzst"): + """Analyze archive contents and compression.""" + contents = list_archive(archive_path, verbose=True) + df = pd.DataFrame(contents) + + # File type analysis + df['extension'] = df['name'].str.split('.').str[-1] + file_types = df.groupby('extension')['size'].agg(['count', 'sum']).reset_index() + + # Visualization + fig, (ax1, ax2) = plt.subplots(1, 2, figsize=(12, 5)) + + # File count by type + ax1.bar(file_types['extension'], file_types['count']) + ax1.set_title('File Count by Type') + ax1.set_xlabel('File Extension') + ax1.set_ylabel('Count') + + # Size by type + ax2.bar(file_types['extension'], file_types['sum'] / 1024) # KB + ax2.set_title('Total Size by Type (KB)') + ax2.set_xlabel('File Extension') + ax2.set_ylabel('Size (KB)') + + plt.tight_layout() + plt.show() + + return df, file_types + +# Execute +contents_df, file_summary = analyze_archive() +print("File Summary:") +display(file_summary) +``` + +These examples demonstrate the flexibility and power of tzst for various real-world scenarios. The library's clean API and robust error handling make it suitable for everything from simple backup scripts to complex enterprise applications. diff --git a/docs/index.md b/docs/index.md index d6ebaab..ef06b58 100644 --- a/docs/index.md +++ b/docs/index.md @@ -14,48 +14,113 @@ README ## What is tzst? -**tzst** is a Python library built exclusively for Python 3.12+ that provides enterprise-grade solutions for handling `.tzst`/`.tar.zst` archives. It combines atomic operations, streaming efficiency, and a meticulously crafted API to redefine how developers handle compressed archives in production environments. +**tzst** is a modern Python library built exclusively for Python 3.12+ that provides comprehensive support for creating, extracting, and managing `.tzst` and `.tar.zst` archives. It combines the proven reliability of the tar format with the superior compression efficiency of Zstandard (zstd) to deliver: + +- **Superior Performance**: Fast compression and decompression with excellent compression ratios +- **Enterprise-Grade Security**: Safe extraction with built-in protections against path traversal attacks +- **Memory Efficiency**: Streaming mode for handling large archives with minimal memory usage +- **Cross-Platform Compatibility**: Works seamlessly on Windows, macOS, and Linux +- **Developer-Friendly**: Clean, Pythonic API with comprehensive error handling ## Key Features -- **🚀 High Performance**: Leverages Zstandard compression for superior speed and compression ratios -- **🔒 Security First**: Built-in extraction filters protect against malicious archives -- **⚡ Streaming Support**: Memory-efficient handling of large archives -- **🛡️ Atomic Operations**: Ensures data integrity with fail-safe file operations -- **🎯 Modern API**: Clean, intuitive interface designed for Python 3.12+ -- **📦 CLI Tools**: Comprehensive command-line interface for everyday tasks +### 🗜️ Advanced Compression + +- **Zstandard Compression**: Best-in-class compression algorithm with configurable levels (1-22) +- **Multiple Extensions**: Support for both `.tzst` and `.tar.zst` file extensions +- **Streaming Support**: Memory-efficient processing for large archives + +### 🔒 Security First + +- **Safe by Default**: Uses 'data' filter for secure extraction without dangerous path traversal +- **Multiple Filter Options**: Choose from 'data', 'tar', or 'fully_trusted' filters based on your security needs +- **Atomic Operations**: All file operations use temporary files with atomic moves to prevent corruption + +### 💻 Dual Interfaces + +- **Command Line**: Intuitive CLI with comprehensive options for batch operations +- **Python API**: Clean, object-oriented interface for programmatic use +- **Convenience Functions**: High-level functions for common operations + +### ⚡ High Performance + +- **Optimized I/O**: Efficient buffering and streaming for large files +- **Conflict Resolution**: Intelligent handling of file conflicts during extraction +- **Cross-Platform**: Native performance on all major operating systems ## Quick Example ```python -from tzst import TzstArchive +from tzst import TzstArchive, create_archive, extract_archive -# Create a new archive -with TzstArchive("backup.tzst", "w", compression_level=5) as archive: - archive.add("documents/") - archive.add("photos/", recursive=True) +# Create an archive +create_archive("backup.tzst", ["documents/", "photos/"], compression_level=5) -# Extract with security -with TzstArchive("backup.tzst", "r") as archive: - archive.extract("documents/", filter="data") +# Extract an archive +extract_archive("backup.tzst", "restore/") + +# Work with archives programmatically +with TzstArchive("data.tzst", "r") as archive: + contents = archive.list(verbose=True) + archive.extract("important.txt", "output/") + is_valid = archive.test() ``` ## Installation -Install tzst from PyPI: +### From PyPI ```bash pip install tzst ``` +### From Source + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install . +``` + +### Standalone Binaries + +Download platform-specific standalone executables from [GitHub Releases](https://github.com/xixu-me/tzst/releases) - no Python installation required! + ## Getting Started -For a quick introduction to using tzst, see the {doc}`quickstart` guide. +For a quick introduction, see the {doc}`quickstart` guide. For comprehensive usage examples, explore the {doc}`examples` section. -For detailed API documentation, browse the {doc}`api/index` section. +### Installation Options + +1. **PyPI Installation** (Recommended): `pip install tzst` +2. **Standalone Binaries**: Download from [GitHub Releases](https://github.com/xixu-me/tzst/releases) +3. **From Source**: Clone and install from repository + +### API Documentation + +Complete API documentation is available in the {doc}`api/index` section, covering: + +- {doc}`api/core`: Main classes and functions +- {doc}`api/cli`: Command-line interface +- {doc}`api/exceptions`: Error handling ## Indices and tables +- :ref:`genindex` +- :ref:`modindex` +- :ref:`search` + +1. **{doc}`quickstart`** - Get up and running quickly with basic examples +2. **{doc}`examples`** - Comprehensive usage examples and patterns +3. **{doc}`api/index`** - Complete API reference documentation + +## Requirements + +- Python 3.12 or higher +- zstandard >= 0.19.0 + +## Reference Links + - {ref}`genindex` - {ref}`modindex` - {ref}`search` diff --git a/docs/quickstart.md b/docs/quickstart.md index 6aff8d6..3da64e9 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -1,222 +1,366 @@ # Quick Start Guide -This guide will help you get started with tzst quickly and efficiently. +This guide will get you up and running with tzst in just a few minutes. ## Installation -Install tzst using pip: +Choose your preferred installation method: + +### Option 1: PyPI (Recommended) ```bash pip install tzst ``` +### Option 2: Standalone Binary + +Download the appropriate executable from [GitHub Releases](https://github.com/xixu-me/tzst/releases): + +| Platform | Architecture | Download | +|----------|--------------|----------| +| **🐧 Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | +| **🐧 Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | +| **🪟 Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | +| **🪟 Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | +| **🍎 macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | +| **🍎 macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | + +Extract the archive and add the executable to your PATH. + +### Option 3: From Source + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install . +``` + ## Basic Usage -### Creating Archives +### Command Line Interface -Use the `TzstArchive` class or convenience functions to create archives: - -```python -from tzst import TzstArchive, create_archive - -# Using TzstArchive class -with TzstArchive("my_archive.tzst", "w", compression_level=5) as archive: - archive.add("file.txt") - archive.add("directory/", recursive=True) - -# Using convenience function -create_archive( - archive_path="backup.tzst", - files=["documents/", "photos/", "config.txt"], - compression_level=10 -) -``` - -### Extracting Archives - -Extract archives safely with built-in security filters: - -```python -from tzst import TzstArchive, extract_archive - -# Using TzstArchive class -with TzstArchive("my_archive.tzst", "r") as archive: - # Extract all files with security filter - archive.extract("output/", filter="data") - - # Extract specific files - archive.extract("output/", members=["file.txt"], filter="data") - -# Using convenience function -extract_archive("backup.tzst", "restore/") -``` - -### Listing Archive Contents - -View what's inside an archive: - -```python -from tzst import TzstArchive, list_archive - -# Using TzstArchive class -with TzstArchive("my_archive.tzst", "r") as archive: - contents = archive.list(verbose=True) - for item in contents: - print(f"{item['name']} - {item['size']} bytes") - -# Using convenience function -files = list_archive("backup.tzst", verbose=True) -``` - -### Testing Archive Integrity - -Verify that an archive is valid: - -```python -from tzst import TzstArchive, test_archive - -# Using TzstArchive class -with TzstArchive("my_archive.tzst", "r") as archive: - is_valid = archive.test() - print(f"Archive is {'valid' if is_valid else 'corrupted'}") - -# Using convenience function -if test_archive("backup.tzst"): - print("Archive is valid") -``` - -## Command Line Interface - -tzst provides a comprehensive CLI for archive operations: - -### Creating Archives +The CLI provides four main operations: ```bash -# Create an archive with multiple files -tzst a backup.tzst documents/ photos/ config.txt +# Create an archive +tzst a archive.tzst file1.txt file2.txt directory/ + +# Extract an archive +tzst x archive.tzst + +# List archive contents +tzst l archive.tzst + +# Test archive integrity +tzst t archive.tzst +``` + +#### Create Archives + +```bash +# Create archive with default compression (level 3) +tzst a backup.tzst documents/ photos/ # Create with high compression -tzst a -l 15 backup.tzst large_files/ +tzst a backup.tzst documents/ photos/ --compression-level 9 -# Create without atomic operations (faster, less safe) -tzst a --no-atomic backup.tzst files/ +# Create from current directory +tzst a project.tzst . + +# Specify different output location +tzst a /backups/data.tzst /home/user/important/ ``` -### Extracting Archives +#### Extract Archives ```bash -# Extract all files (default: safe extraction) +# Extract to current directory tzst x backup.tzst # Extract to specific directory -tzst x backup.tzst -o restore/ +tzst x backup.tzst --output /restore/ # Extract specific files only -tzst x backup.tzst config.txt documents/ +tzst x backup.tzst documents/report.pdf photos/vacation.jpg -# Extract with streaming (memory efficient) -tzst x backup.tzst --streaming +# Extract with conflict resolution +tzst x backup.tzst --conflict-resolution skip ``` -### Listing Contents +#### List Contents ```bash # Simple listing tzst l backup.tzst # Detailed listing with file info -tzst l backup.tzst -v +tzst l backup.tzst --verbose -# Streaming mode for large archives -tzst l backup.tzst --streaming +# Stream large archives efficiently +tzst l huge-archive.tzst --streaming ``` -### Testing Archives +### Python API -```bash -# Test archive integrity -tzst t backup.tzst - -# Test with streaming -tzst t backup.tzst --streaming -``` - -## Security Considerations - -tzst includes built-in security features to protect against malicious archives: - -### Extraction Filters - -Always use appropriate filters when extracting archives from untrusted sources: - -- **`data`** (default): Safest option, only extracts regular files and directories -- **`tar`**: Honors most tar features but still secure -- **`fully_trusted`**: No restrictions (only use with completely trusted archives) +#### Quick Start ```python -# Safe extraction (recommended) -archive.extract("output/", filter="data") +from tzst import create_archive, extract_archive, list_archive, test_archive -# Command line -tzst x archive.tzst --filter=data +# Create an archive +create_archive("backup.tzst", ["documents/", "photos/"], compression_level=5) + +# Extract an archive +extract_archive("backup.tzst", "restore/") + +# List contents +contents = list_archive("backup.tzst", verbose=True) +for item in contents: + print(f"{item['name']} - {item['size']} bytes") + +# Test integrity +is_valid = test_archive("backup.tzst") +print(f"Archive is {'valid' if is_valid else 'corrupted'}") ``` -### Best Practices - -1. **Always use the default `data` filter** for untrusted archives -2. **Enable atomic operations** (default) for data integrity -3. **Use streaming mode** for very large archives to save memory -4. **Validate archives** with `test()` before processing -5. **Specify output directories** explicitly to avoid overwrites - -## Performance Tips - -### Memory Efficiency - -For large archives, use streaming mode: +#### Using the TzstArchive Class ```python -# Streaming mode uses less memory -with TzstArchive("large.tzst", "r", streaming=True) as archive: - archive.extract("output/") +from tzst import TzstArchive + +# Create a new archive +with TzstArchive("data.tzst", "w", compression_level=6) as archive: + archive.add("file.txt") + archive.add("directory/", recursive=True) + + # Add with custom archive name + archive.add("config/prod.yaml", arcname="config.yaml") + +# Read an existing archive +with TzstArchive("data.tzst", "r") as archive: + # List contents + contents = archive.list(verbose=True) + for item in contents: + print(f"{item['name']} - {item['size']} bytes") + + # Test integrity + is_valid = archive.test() + print(f"Archive is {'valid' if is_valid else 'corrupted'}") + + # Extract specific files + archive.extract("file.txt", "output/") + + # Extract all files + archive.extractall("restore/") ``` -### Compression Levels +## Advanced Features -Choose appropriate compression levels based on your needs: - -- **Level 1-3**: Fast compression, larger files -- **Level 3-6**: Balanced (default: 3) -- **Level 7-15**: Better compression, slower -- **Level 16-22**: Maximum compression, much slower +### Security and Filtering ```python -# Fast compression for temporary files -TzstArchive("temp.tzst", "w", compression_level=1) +from tzst import extract_archive -# Maximum compression for long-term storage -TzstArchive("backup.tzst", "w", compression_level=15) +# Safe extraction with built-in security (default) +extract_archive("untrusted.tzst", "safe-output/", filter="data") + +# For trusted archives with special features +extract_archive("trusted.tzst", "output/", filter="tar") +``` + +### Conflict Resolution + +```python +from tzst import extract_archive, ConflictResolution + +# Skip existing files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.SKIP_ALL) + +# Auto-rename conflicting files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.AUTO_RENAME_ALL) +``` + +### Performance Optimization + +```python +from tzst import create_archive, extract_archive + +# Create with different compression levels +create_archive("fast.tzst", files, compression_level=1) # Fastest +create_archive("balanced.tzst", files, compression_level=6) # Balanced +create_archive("best.tzst", files, compression_level=22) # Best compression + +# Memory-efficient operations for large archives +extract_archive("huge-archive.tzst", "output/", streaming=True) ``` ## Error Handling -tzst provides specific exceptions for different error conditions: - ```python -from tzst import TzstArchive -from tzst.exceptions import TzstArchiveError, TzstDecompressionError +from tzst import create_archive, TzstArchiveError, TzstCompressionError try: - with TzstArchive("archive.tzst", "r") as archive: - archive.extract("output/") + create_archive("backup.tzst", ["documents/"]) +except TzstCompressionError as e: + print(f"Compression failed: {e}") except TzstArchiveError as e: - print(f"Archive error: {e}") -except TzstDecompressionError as e: - print(f"Decompression error: {e}") + print(f"Archive operation failed: {e}") +except Exception as e: + print(f"Unexpected error: {e}") ``` ## Next Steps -- Explore the complete {doc}`api/index` documentation -- Check out more {doc}`examples` and use cases -- Read about advanced features in the full documentation +- Explore comprehensive {doc}`examples` for real-world scenarios +- Check the {doc}`api/index` for detailed API documentation +- See advanced features like atomic operations and custom filters +- Learn about integration with web frameworks and automation tools + +## Read an Existing Archive + +```python +with TzstArchive("data.tzst", "r") as archive: + # List contents + contents = archive.list(verbose=True) + + # Extract specific file + archive.extract("file.txt", "output/") + + # Test integrity + is_valid = archive.test() + + # Get raw member information + members = archive.getmembers() +``` + +## Important Concepts + +### Compression Levels + +tzst supports compression levels from 1 to 22: + +- **Level 1-3**: Fast compression, larger files (good for temporary archives) +- **Level 4-6**: Balanced compression and speed (recommended for most use cases) +- **Level 7-15**: Higher compression, slower (good for long-term storage) +- **Level 16-22**: Maximum compression, much slower (for size-critical applications) + +```python +# Fast compression +create_archive("temp.tzst", files, compression_level=1) + +# Balanced (default) +create_archive("backup.tzst", files, compression_level=3) + +# High compression +create_archive("archive.tzst", files, compression_level=9) + +# Maximum compression +create_archive("minimal.tzst", files, compression_level=22) +``` + +### Security Filters + +tzst provides extraction filters to protect against malicious archives: + +```python +# Safe data extraction (default, recommended) +extract_archive("archive.tzst", "output/", filter="data") + +# Preserve more tar features but still secure +extract_archive("archive.tzst", "output/", filter="tar") + +# Full trust mode (use only with trusted archives) +extract_archive("archive.tzst", "output/", filter="fully_trusted") +``` + +### Streaming Mode + +For large archives (>100MB), use streaming mode to reduce memory usage: + +```python +# Memory-efficient operations +with TzstArchive("large-archive.tzst", "r", streaming=True) as archive: + contents = archive.list() + archive.extractall("output/") + is_valid = archive.test() +``` + +**Note**: Streaming mode has limitations - you cannot extract specific files or use random access operations. + +### Handling File Conflicts + +Handle file conflicts during extraction: + +```python +from tzst import ConflictResolution + +# Skip existing files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.SKIP) + +# Replace all existing files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.REPLACE_ALL) + +# Auto-rename conflicting files +extract_archive("archive.tzst", "output/", + conflict_resolution=ConflictResolution.AUTO_RENAME_ALL) +``` + +## Common Patterns + +### Backup Script + +```python +#!/usr/bin/env python3 +from pathlib import Path +from datetime import datetime +from tzst import create_archive + +def create_backup(): + timestamp = datetime.now().strftime("%Y%m%d_%H%M%S") + backup_name = f"backup_{timestamp}.tzst" + + # Backup important directories + directories = ["documents/", "projects/", "config/"] + + print(f"Creating backup: {backup_name}") + create_archive(backup_name, directories, compression_level=6) + print(f"Backup created: {Path(backup_name).stat().st_size / 1024 / 1024:.1f} MB") + +if __name__ == "__main__": + create_backup() +``` + +### Archive Verification + +```python +from tzst import test_archive, list_archive + +def verify_archive(archive_path): + print(f"Verifying {archive_path}...") + + # Test integrity + if not test_archive(archive_path): + print("❌ Archive is corrupted!") + return False + + # List contents + contents = list_archive(archive_path, verbose=True) + total_size = sum(item['size'] for item in contents if item['is_file']) + file_count = sum(1 for item in contents if item['is_file']) + + print(f"✅ Archive is valid") + print(f"📁 Files: {file_count}") + print(f"📦 Total size: {total_size / 1024 / 1024:.1f} MB") + + return True +``` + +## Further Learning + +- Explore {doc}`examples` for more advanced usage patterns +- Check the {doc}`api/index` for complete API documentation +- Read the full {doc}`README` for additional features and background