diff --git a/docs/404.md b/docs/404.md index afcbf70..261c6f3 100644 --- a/docs/404.md +++ b/docs/404.md @@ -11,7 +11,7 @@ myst: # 404 - Page Not Found -## 🔍 Oops! The page you're looking for doesn't exist +## Oops! The page you're looking for doesn't exist The URL you requested could not be found in the tzst documentation. This might happen if: @@ -20,14 +20,17 @@ The URL you requested could not be found in the tzst documentation. This might h - There's a typo in the URL - The page has been removed -## 🏠 Where would you like to go? +## Where would you like to go? ### Popular Pages - **{doc}`index`** - Documentation homepage - **{doc}`quickstart`** - Get started with tzst +- **{doc}`performance`** - Learn about performance optimizations - **{doc}`examples`** - Practical examples and use cases - **{doc}`api/index`** - Complete API reference +- **{doc}`development`** - Contributing to tzst +- **{ref}`genindex`** - Index of all documented items ### Quick Navigation @@ -36,7 +39,7 @@ The URL you requested could not be found in the tzst documentation. This might h - **Security Features** - {ref}`security-and-filtering` - **Error Handling** - {ref}`error-handling` -## 🔎 Search Documentation +## Search Documentation Use the search box in the top navigation to find what you're looking for, or browse through these sections: @@ -53,13 +56,13 @@ Use the search box in the top navigation to find what you're looking for, or bro - **Integration Examples** - {ref}`integration-examples` - **Performance Optimization** - {ref}`performance-optimization` -## 📚 Additional Resources +## Additional Resources - [GitHub Repository](https://github.com/xixu-me/tzst) - Source code and issue tracker - [PyPI Package](https://pypi.org/project/tzst/) - Download and installation - [Release Notes](https://github.com/xixu-me/tzst/releases) - Latest updates and changes -## 🐛 Report an Issue +## Report an Issue If you believe this is a broken link within our documentation, please [report it on GitHub](https://github.com/xixu-me/tzst/issues). diff --git a/docs/api/cli.md b/docs/api/cli.md index 7f01546..e07dcc0 100644 --- a/docs/api/cli.md +++ b/docs/api/cli.md @@ -25,17 +25,17 @@ The command-line interface module provides comprehensive functionality for the t The tzst CLI provides a powerful command-line interface for archive operations with intuitive commands and comprehensive options. The interface is designed for both interactive use and scripting, with robust error handling and user-friendly output. -### 🔧 Core Commands +### Core Commands | Command | Aliases | Description | Streaming Support | |---------|---------|-------------|-------------------| | `a` | `add`, `create` | Create or add to archive | N/A | -| `x` | `extract` | Extract with full paths | ✓ `--streaming` | -| `e` | `extract-flat` | Extract without directory structure | ✓ `--streaming` | -| `l` | `list` | List archive contents | ✓ `--streaming` | -| `t` | `test` | Test archive integrity | ✓ `--streaming` | +| `x` | `extract` | Extract with full paths | `--streaming` | +| `e` | `extract-flat` | Extract without directory structure | `--streaming` | +| `l` | `list` | List archive contents | `--streaming` | +| `t` | `test` | Test archive integrity | `--streaming` | -### 🎯 Key Features +### Key Features - **Intuitive Commands**: Simple, memorable command aliases (a, x, e, l, t) - **Streaming Support**: Memory-efficient processing for large archives diff --git a/docs/api/index.md b/docs/api/index.md index a3c11b8..308a71a 100644 --- a/docs/api/index.md +++ b/docs/api/index.md @@ -104,26 +104,26 @@ Exception hierarchy for comprehensive error handling and debugging support. ## Key Features -### 🛡️ Security First +### Security First - Built-in path traversal protection - Multiple security filter options - Safe extraction by default -### ⚡ High Performance +### High Performance - Zstandard compression with configurable levels - Streaming support for large archives - Memory-efficient operations -### 🔧 Developer Friendly +### Developer Friendly - Clean, Pythonic API - Comprehensive error handling - Context manager support - Extensive documentation and examples -### 🌐 Cross-Platform +### Cross-Platform - Works on Windows, macOS, and Linux - Consistent behavior across platforms diff --git a/docs/development.md b/docs/development.md new file mode 100644 index 0000000..10f0556 --- /dev/null +++ b/docs/development.md @@ -0,0 +1,353 @@ +# Development Guide + +This guide provides comprehensive information for developers contributing to or working with the tzst library. + +## Setting up Development Environment + +This project uses modern Python packaging standards: + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install -e .[dev] +``` + +The development installation includes all necessary tools: + +- **pytest** - Testing framework +- **ruff** - Linting and formatting +- **coverage** - Code coverage analysis +- **sphinx** - Documentation generation + +## Running Tests + +### Basic Test Commands + +```bash +# Run all tests +python -m pytest + +# Run tests with coverage +pytest --cov=tzst --cov-report=html + +# Or use the simpler command (coverage settings are in pyproject.toml) +pytest + +# Run with verbose output +python -m pytest -v + +# Run specific test file +python -m pytest tests/test_core.py + +# Run integration tests only +python -m pytest -m integration +``` + +### Test Structure + +- **Unit tests**: Test individual functions and methods +- **Integration tests**: Test component interactions +- **CLI tests**: Test command-line interface +- **Platform-specific tests**: Test OS-specific functionality + +### Writing Tests + +1. **Use descriptive test names:** + + ```python + def test_create_archive_with_compression_level_9(): + ``` + +2. **Use fixtures for common test data:** + + ```python + def test_extract_archive(sample_archive_path, temp_dir): + ``` + +3. **Test edge cases:** + - Empty files + - Large files + - Invalid inputs + - Corrupted archives + +4. **Add markers for test categorization:** + + ```python + @pytest.mark.integration + def test_full_archive_workflow(): + ``` + +## Code Quality + +### Running Code Style Tools + +```bash +# Check code quality +ruff check src tests + +# Fix auto-fixable issues +ruff check --fix src tests + +# Format code +ruff format src tests + +# Check formatting without making changes +ruff format --check src tests +``` + +### Configuration + +Settings are defined in `pyproject.toml`: + +- Line length: 88 characters +- Target Python version: 3.12+ +- Import sorting with isort +- Quote style: double quotes + +### Code Style Guidelines + +1. **Follow PEP 8** with project-specific modifications +2. **Use type hints** for all public APIs +3. **Write docstrings** for classes and public methods +4. **Keep functions focused** and reasonably sized +5. **Use meaningful variable names** +6. **Add comments** for complex logic + +## Documentation + +### Building Documentation + +```bash +# Navigate to docs directory +cd docs + +# Install documentation dependencies +pip install -r requirements.txt + +# Build HTML documentation +make html + +# On Windows, use: +make.bat html + +# View built documentation +# Open docs/_build/html/index.html in your browser +``` + +### Documentation Structure + +``` +docs/ +├── index.md # Main documentation landing page +├── quickstart.md # Getting started guide +├── performance.md # Performance guide and comparisons +├── examples.md # Usage examples +├── development.md # This development guide +├── api/ # API reference documentation +│ ├── index.md +│ ├── core.md +│ ├── cli.md +│ └── exceptions.md +├── conf.py # Sphinx configuration +└── requirements.txt # Documentation dependencies +``` + +### Writing Documentation + +- Use **MyST Markdown** format +- Include **code examples** for new features +- Add **cross-references** using proper syntax +- Test all **code snippets** to ensure they work + +## Project Structure + +``` +tzst/ +├── src/tzst/ # Main package source code +│ ├── __init__.py # Package initialization and exports +│ ├── __main__.py # CLI entry point +│ ├── cli.py # Command-line interface +│ ├── core.py # Core archive functionality +│ └── exceptions.py # Custom exceptions +├── tests/ # Test suite +│ ├── conftest.py # Pytest configuration and fixtures +│ ├── test_core.py # Core functionality tests +│ ├── test_cli.py # CLI tests +│ └── test_*.py # Additional test modules +├── docs/ # Documentation source +├── .github/ # GitHub workflows and templates +├── pyproject.toml # Project configuration +├── README.md # Project documentation +├── LICENSE # BSD 3-Clause License +└── CONTRIBUTING.md # Contribution guidelines +``` + +## 🤝 Contributing Workflow + +### 1. Making Changes + +#### Types of Contributions + +- **Bug fixes**: Fix issues in existing functionality +- **Features**: Add new capabilities to the library +- **Documentation**: Improve or add documentation +- **Tests**: Add or improve test coverage +- **Performance**: Optimize existing code +- **Security**: Address security vulnerabilities + +#### Branch Naming + +Use descriptive branch names: + +- `feature/add-streaming-mode` +- `fix/handle-corrupted-archives` +- `docs/improve-api-documentation` +- `test/add-compression-tests` + +### 2. Commit Messages + +Follow conventional commit format: + +``` +type(scope): description + +[optional body] + +[optional footer] +``` + +**Types:** + +- `feat`: New feature +- `fix`: Bug fix +- `docs`: Documentation changes +- `test`: Adding or modifying tests +- `refactor`: Code refactoring +- `perf`: Performance improvements +- `chore`: Build process or auxiliary tool changes + +**Examples:** + +``` +feat(core): add streaming compression support + +fix(cli): handle invalid archive paths gracefully + +docs(readme): update installation instructions +``` + +### 3. Pull Request Process + +1. **Create a feature branch:** + + ```bash + git checkout -b feature/your-feature-name + ``` + +2. **Make your changes** following the guidelines above + +3. **Add tests** for new functionality + +4. **Update documentation** if needed + +5. **Run the test suite:** + + ```bash + python -m pytest + ruff check . + ruff format --check . + ``` + +6. **Commit your changes:** + + ```bash + git add . + git commit -m "feat: add your feature description" + ``` + +7. **Push to your fork:** + + ```bash + git push origin feature/your-feature-name + ``` + +8. **Create a pull request** using the provided template + +### 4. Pull Request Guidelines + +- **Fill out the PR template** completely +- **Link related issues** using keywords (fixes #123) +- **Keep PRs focused** - one feature/fix per PR +- **Ensure all CI checks pass** +- **Respond to review feedback** promptly + +## Development Tips + +### Performance Considerations + +- Use streaming for large files +- Consider memory usage patterns +- Profile code for bottlenecks +- Test with various file sizes + +### Security Considerations + +- Validate all user inputs +- Use secure defaults (e.g., 'data' filter) +- Handle malicious archives safely +- Be cautious with file paths + +### Compatibility + +- Support Python 3.12+ +- Test on multiple platforms (Windows, macOS, Linux) +- Consider different filesystem behaviors +- Maintain backwards compatibility when possible + +## Release Process + +Releases are handled by maintainers: + +1. Update version in `src/tzst/__init__.py` +2. Create a release tag +3. Automated CI/CD publishes to PyPI + +## Getting Help + +### Resources + +- **Issues**: [GitHub Issues](https://github.com/xixu-me/tzst/issues) +- **Discussions**: Use GitHub Discussions for questions +- **Documentation**: Check the README and code comments + +### Reporting Issues + +When reporting bugs: + +1. **Use the bug report template** +2. **Provide a minimal reproduction case** +3. **Include system information** (OS, Python version) +4. **Attach relevant files** if possible (archives, logs) + +### Suggesting Features + +When suggesting features: + +1. **Use the feature request template** +2. **Explain the use case** and motivation +3. **Consider backwards compatibility** +4. **Provide implementation ideas** if you have them + +## Code of Conduct + +This project follows the principles of respectful collaboration. Please be kind, constructive, and professional in all interactions. + +## Recognition + +Contributors are recognized in several ways: + +- Listed in release notes for significant contributions +- Mentioned in README acknowledgments +- GitHub contributor statistics + +Thank you for contributing to tzst! Your efforts help make this library better for everyone. diff --git a/docs/examples.md b/docs/examples.md index 1a47103..c1052d8 100644 --- a/docs/examples.md +++ b/docs/examples.md @@ -89,14 +89,19 @@ for item in contents: ## Command Line Usage +> **Note**: Download the [standalone binary](https://github.com/xixu-me/tzst/releases) for the best performance and no Python dependency. Alternatively, use `uvx tzst` for running without installation. See [uv documentation](https://docs.astral.sh/uv/) for details. + ### Archive Creation Commands ```bash # Basic archive creation tzst a backup.tzst documents/ photos/ +# Or with uvx (no installation needed) +uvx tzst a backup.tzst documents/ photos/ # Create with specific compression level tzst a backup.tzst documents/ photos/ -l 6 +uvx tzst a backup.tzst documents/ photos/ -l 6 # Create from multiple sources tzst a complete-backup.tzst /home/user/documents /home/user/photos /etc/config @@ -110,6 +115,8 @@ tzst a backup.tzst documents/ photos/ -v ```bash # Extract to current directory tzst x backup.tzst +# Or with uvx +uvx tzst x backup.tzst # Extract to specific directory tzst x backup.tzst --output /restore/ @@ -129,6 +136,8 @@ tzst e backup.tzst --output flat-restore/ ```bash # List archive contents tzst l backup.tzst +# Or with uvx +uvx tzst l backup.tzst # List with detailed information tzst l backup.tzst --verbose @@ -551,23 +560,23 @@ def validate_and_repair_archive(archive_path): # Test basic integrity try: if test_archive(archive_path): - print("✅ Archive integrity test passed") + print("Archive integrity test passed") return True except Exception as e: - print(f"❌ Integrity test failed: {e}") + print(f"Integrity test failed: {e}") # Try to list contents try: contents = list_archive(archive_path) - print(f"📁 Archive contains {len(contents)} items") + print(f"Archive contains {len(contents)} items") # Try streaming mode if regular mode fails contents_streaming = list_archive(archive_path, streaming=True) if len(contents_streaming) != len(contents): - print("⚠️ Different results between modes - possible corruption") + print("Different results between modes - possible corruption") except Exception as e: - print(f"❌ Cannot list contents: {e}") + print(f"Cannot list contents: {e}") return False # Try partial extraction @@ -583,13 +592,13 @@ def validate_and_repair_archive(archive_path): archive.extract(member.name, backup_dir) extracted_count += 1 except Exception as e: - print(f"⚠️ Failed to extract {member.name}: {e}") + print(f"Failed to extract {member.name}: {e}") - print(f"✅ Recovered {extracted_count} files to {backup_dir}") + print(f"Recovered {extracted_count} files to {backup_dir}") return True except Exception as e: - print(f"❌ Recovery failed: {e}") + print(f"Recovery failed: {e}") return False # Example usage @@ -648,15 +657,15 @@ class BackupManager: # Validate the backup if test_archive(backup_path): file_size = backup_path.stat().st_size / (1024 * 1024) - print(f"✅ Backup created and validated: {file_size:.1f} MB") + print(f"Backup created and validated: {file_size:.1f} MB") return backup_path else: - print("❌ Backup validation failed!") + print("Backup validation failed!") backup_path.unlink() # Remove invalid backup return None except Exception as e: - print(f"❌ Backup failed: {e}") + print(f"Backup failed: {e}") return None def cleanup_old_backups(self): @@ -792,13 +801,13 @@ def archive_logs_by_date(log_dir, archive_dir, days_old=7): archived_files.append(log_file) file_size = archive_path.stat().st_size / 1024 - print(f"✅ Created {archive_name} ({file_size:.1f} KB)") + print(f"Created {archive_name} ({file_size:.1f} KB)") else: - print(f"❌ Archive validation failed for {archive_name}") + print(f"Archive validation failed for {archive_name}") archive_path.unlink() except Exception as e: - print(f"❌ Failed to archive logs for {date_key}: {e}") + print(f"Failed to archive logs for {date_key}: {e}") print(f"Archived {len(archived_files)} log files") @@ -856,9 +865,8 @@ class DataMigrator: with open(checksum_file, "w") as f: f.write(f"{checksum} {package_path.name}\n") - package_size = package_path.stat().st_size / (1024 * 1024) - print(f"✅ Package created: {package_size:.1f} MB") - print(f"📋 Checksum: {checksum}") + package_size = package_path.stat().st_size / (1024 * 1024) print(f"Package created: {package_size:.1f} MB") + print(f"Checksum: {checksum}") return package_path, checksum_file @@ -877,13 +885,13 @@ class DataMigrator: if expected_checksum != actual_checksum: raise RuntimeError(f"Checksum mismatch! Expected: {expected_checksum}, Got: {actual_checksum}") - print("✅ Checksum verification passed") + print("Checksum verification passed") # Test archive integrity if not test_archive(package_path): raise RuntimeError("Archive integrity check failed") - print("✅ Archive integrity verified") + print("Archive integrity verified") # Extract with conflict resolution destination.mkdir(parents=True, exist_ok=True) @@ -893,7 +901,7 @@ class DataMigrator: conflict_resolution="replace_all" # Overwrite for migration ) - print(f"✅ Package extracted to: {destination}") + print(f"Package extracted to: {destination}") # Example usage if __name__ == "__main__": diff --git a/docs/index.md b/docs/index.md index 13070b4..23acc21 100644 --- a/docs/index.md +++ b/docs/index.md @@ -11,6 +11,13 @@ myst: # tzst Documentation +[![codecov](https://codecov.io/gh/xixu-me/tzst/graph/badge.svg?token=2AIN1559WU)](https://codecov.io/gh/xixu-me/tzst) +[![CodeQL](https://github.com/xixu-me/tzst/actions/workflows/github-code-scanning/codeql/badge.svg)](https://github.com/xixu-me/tzst/actions/workflows/github-code-scanning/codeql) +[![CI/CD](https://github.com/xixu-me/tzst/actions/workflows/ci.yml/badge.svg)](https://github.com/xixu-me/tzst/actions/workflows/ci.yml) +[![PyPI - Version](https://img.shields.io/pypi/v/tzst)](https://pypi.org/project/tzst/) +[![GitHub License](https://img.shields.io/github/license/xixu-me/tzst)](LICENSE) +[![Sponsor](https://img.shields.io/badge/Sponsor-violet)](https://xi-xu.me/#sponsorships) + Welcome to **tzst**, the next-generation Python library engineered for modern archive management, leveraging cutting-edge Zstandard compression to deliver superior performance, security, and reliability. ```{toctree} @@ -18,8 +25,11 @@ Welcome to **tzst**, the next-generation Python library engineered for modern ar :caption: Contents: quickstart -api/index +performance examples +api/index +development +genindex ``` ```{toctree} @@ -41,25 +51,25 @@ README ## Key Features -### 🗜️ Advanced Compression +### Advanced Compression - **Zstandard Compression**: Best-in-class compression algorithm with configurable levels (1-22) - **Multiple Extensions**: Support for both `.tzst` and `.tar.zst` file extensions - **Streaming Support**: Memory-efficient processing for large archives -### 🔒 Security First +### Security First - **Safe by Default**: Uses 'data' filter for secure extraction without dangerous path traversal - **Multiple Filter Options**: Choose from 'data', 'tar', or 'fully_trusted' filters based on your security needs - **Atomic Operations**: All file operations use temporary files with atomic moves to prevent corruption -### 💻 Dual Interfaces +### Dual Interfaces - **Command Line**: Intuitive CLI with comprehensive options for batch operations - **Python API**: Clean, object-oriented interface for programmatic use - **Convenience Functions**: High-level functions for common operations -### ⚡ High Performance +### High Performance - **Optimized I/O**: Efficient buffering and streaming for large files - **Conflict Resolution**: Intelligent handling of file conflicts during extraction @@ -91,6 +101,49 @@ with TzstArchive("data.tzst", "r") as archive: pip install tzst ``` +### From GitHub Releases + +Download platform-specific standalone executables from [GitHub Releases](https://github.com/xixu-me/tzst/releases) - no Python installation required! + +#### Supported Platforms + +| Platform | Architecture | File | +|----------|-------------|------| +| **Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | +| **Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | +| **Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | +| **Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | +| **macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | +| **macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | + +#### Installation Steps + +1. **Download** the appropriate archive for your platform from the [latest releases page](https://github.com/xixu-me/tzst/releases/latest) +2. **Extract** the archive to get the `tzst` executable (or `tzst.exe` on Windows) +3. **Move** the executable to a directory in your PATH: + - **Linux/macOS**: `sudo mv tzst /usr/local/bin/` + - **Windows**: Add the directory containing `tzst.exe` to your PATH environment variable +4. **Verify** installation: `tzst --help` + +#### Benefits of Binary Installation + +- **No Python required** - Standalone executable +- **Faster startup** - No Python interpreter overhead +- **Easy deployment** - Single file distribution +- **Consistent behavior** - Bundled dependencies + +### Using uvx (No Installation) + +Run tzst directly without installation using [uvx](https://docs.astral.sh/uv/): + +```bash +uvx tzst --help +uvx tzst a archive.tzst file1.txt file2.txt directory/ +uvx tzst x archive.tzst +``` + +Perfect for one-time usage, testing, CI/CD pipelines, and isolated environments. + ### From Source ```bash @@ -99,19 +152,16 @@ cd tzst pip install . ``` -### Standalone Binaries - -Download platform-specific standalone executables from [GitHub Releases](https://github.com/xixu-me/tzst/releases) - no Python installation required! - ## Getting Started For a quick introduction, see the {doc}`quickstart` guide. For comprehensive usage examples, explore the {doc}`examples` section. ### Installation Options -1. **PyPI Installation** (Recommended): `pip install tzst` +1. **PyPI Installation**: `pip install tzst` 2. **Standalone Binaries**: Download from [GitHub Releases](https://github.com/xixu-me/tzst/releases) -3. **From Source**: Clone and install from repository +3. **uvx (No Installation)**: Run directly with `uvx tzst` +4. **From Source**: Clone and install from repository ### API Documentation @@ -121,12 +171,63 @@ Complete API documentation is available in the {doc}`api/index` section, coverin - {doc}`api/cli`: Command-line interface - {doc}`api/exceptions`: Error handling -## Indices and tables +## Development + +For comprehensive development information, see the {doc}`development` guide, which covers: + +- Setting up development environment +- Running tests and code quality checks +- Documentation building +- Contributing workflow and guidelines +- Project structure and best practices + +### Quick Start + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install -e .[dev] +pytest +``` + +## Contributing + +We welcome contributions! Please read our [Contributing Guide](https://github.com/xixu-me/tzst/blob/main/CONTRIBUTING.md) for: + +- Development setup and project structure +- Code style guidelines and best practices +- Testing requirements and writing tests +- Pull request process and review workflow + +### Types of Contributions Welcome + +- **Bug fixes** - Fix issues in existing functionality +- **Features** - Add new capabilities to the library +- **Documentation** - Improve or add documentation +- **Tests** - Add or improve test coverage +- **Performance** - Optimize existing code +- **Security** - Address security vulnerabilities + +## Acknowledgments + +- [Meta Zstandard](https://github.com/facebook/zstd) for the excellent compression algorithm +- [python-zstandard](https://github.com/indygreg/python-zstandard) for Python bindings +- The Python community for inspiration and feedback + +## License + +Copyright © 2025 [Xi Xu](https://xi-xu.me). All rights reserved. + +Licensed under the [BSD 3-Clause](https://github.com/xixu-me/tzst/blob/main/LICENSE) license. + +## Documentation Guide 1. **{doc}`quickstart`** - Get up and running quickly with basic examples -2. **{doc}`examples`** - Comprehensive usage examples and patterns -3. **{doc}`api/index`** - Complete API reference documentation -4. **{ref}`genindex`** - Index of all documented items +2. **{doc}`performance`** - Performance optimization guide and comparisons +3. **{doc}`examples`** - Comprehensive usage examples and patterns +4. **{doc}`api/index`** - Complete API reference documentation +5. **{doc}`development`** - Development and contribution guidelines +6. **{ref}`genindex`** - Index of all documented items ## Requirements diff --git a/docs/performance.md b/docs/performance.md new file mode 100644 index 0000000..6677326 --- /dev/null +++ b/docs/performance.md @@ -0,0 +1,278 @@ +--- +myst: + html_meta: + description: "tzst Performance Guide - Compression level optimization, performance tips, and comparison with other archive tools" + keywords: "tzst performance, compression benchmarks, tar gzip comparison, archive performance optimization" + og:title: "tzst Performance Guide" + og:description: "Performance optimization tips and comparison with other archive tools for tzst" + twitter:title: "tzst Performance Guide" + twitter:description: "Performance optimization tips and comparison with other archive tools for tzst" +--- + +# Performance Guide + +This guide covers performance optimization techniques and provides detailed comparisons with other archive tools. + +## Performance Tips + +### 1. Compression Levels + +Choose the right compression level for your use case: + +- **Level 1-3**: Fast compression, larger files (good for temporary archives or real-time processing) +- **Level 3** (default): Optimal balance for most use cases +- **Level 6-9**: Higher compression, moderate speed (good for regular backups) +- **Level 15-22**: Maximum compression, slower (for long-term storage or bandwidth-limited scenarios) + +```python +from tzst import create_archive + +# For temporary files or frequent operations +create_archive("temp.tzst", files, compression_level=1) + +# Balanced default (recommended) +create_archive("backup.tzst", files, compression_level=3) + +# Long-term storage +create_archive("archive.tzst", files, compression_level=9) + +# Maximum compression for critical space savings +create_archive("minimal.tzst", files, compression_level=22) +``` + +### 2. Streaming + +Use streaming mode for archives larger than 100MB: + +```python +from tzst import extract_archive, list_archive, test_archive + +# Memory-efficient operations for large archives +extract_archive("large-backup.tzst", "restore/", streaming=True) +contents = list_archive("large-backup.tzst", streaming=True) +is_valid = test_archive("large-backup.tzst", streaming=True) +``` + +**Streaming Benefits:** + +- Significantly reduced memory usage +- Better performance for large archives +- Handles archives that don't fit in memory + +### 3. Batch Operations + +Add multiple files in a single session when possible: + +```python +from tzst import TzstArchive + +# Efficient: Single archive session +with TzstArchive("backup.tzst", "w") as archive: + archive.add("file1.txt") + archive.add("file2.txt") + archive.add("directory/", recursive=True) + +# Less efficient: Multiple separate operations +create_archive("backup1.tzst", ["file1.txt"]) +create_archive("backup2.tzst", ["file2.txt"]) +``` + +### 4. File Type Considerations + +- Already compressed files (`.jpg`, `.png`, `.mp4`, `.pdf`) won't compress much further +- Text files, source code, and logs compress very well +- Consider compression level based on your data types + +## Comparison with Other Tools + +### vs tar + gzip + +**tzst Advantages:** + +- **Better compression ratios**: 10-40% smaller archives +- **Faster decompression**: 2-3x faster extraction +- **Modern algorithm**: Better handling of various file types +- **Streaming support**: Better memory efficiency + +**When to use tar + gzip:** + +- Legacy system compatibility requirements +- Very old systems without zstd support + +### vs tar + xz + +**tzst Advantages:** + +- **Significantly faster compression**: 3-10x faster creation +- **Faster decompression**: 2-4x faster extraction +- **Better speed/compression trade-off**: Similar compression with much better speed +- **More compression levels**: Fine-grained control (22 levels vs 9) + +**When to use tar + xz:** + +- Maximum compression is critical and time is not a factor +- Systems that don't support zstd + +### vs zip + +**tzst Advantages:** + +- **Better compression**: 15-30% smaller archives +- **Preserves Unix permissions and metadata**: Full POSIX compatibility +- **Better streaming support**: Memory-efficient for large archives +- **Better directory handling**: Preserves directory structure and timestamps + +**When to use zip:** + +- Cross-platform compatibility with very old systems +- Individual file access without full extraction is required +- Windows-centric environments with no command-line tools + +## Benchmarking Examples + +### Compression Level Benchmark + +```python +import time +from pathlib import Path +from tzst import create_archive + +def benchmark_compression_levels(files, output_prefix="benchmark"): + """Compare different compression levels.""" + levels_to_test = [1, 3, 6, 9, 15, 22] + + results = [] + for level in levels_to_test: + output_file = f"{output_prefix}_level_{level}.tzst" + + # Measure compression time + start_time = time.time() + create_archive(output_file, files, compression_level=level) + compress_time = time.time() - start_time + + # Get file size + file_size = Path(output_file).stat().st_size + + results.append({ + 'level': level, + 'time': compress_time, + 'size': file_size, + 'size_mb': file_size / (1024 * 1024) + }) + + print(f"Level {level}: {compress_time:.2f}s, {file_size/1024/1024:.1f} MB") + + return results + +# Example usage +files = ["documents/", "projects/"] +results = benchmark_compression_levels(files) +``` + +### Memory Usage Comparison + +```python +import psutil +import os +from tzst import extract_archive + +def monitor_memory_usage(func, *args, **kwargs): + """Monitor memory usage during function execution.""" + process = psutil.Process(os.getpid()) + initial_memory = process.memory_info().rss / 1024 / 1024 # MB + + func(*args, **kwargs) + + peak_memory = process.memory_info().rss / 1024 / 1024 # MB + return peak_memory - initial_memory + +# Compare streaming vs non-streaming extraction +large_archive = "large-dataset.tzst" + +memory_normal = monitor_memory_usage(extract_archive, large_archive, "output1/") +memory_streaming = monitor_memory_usage(extract_archive, large_archive, "output2/", streaming=True) + +print(f"Normal extraction: {memory_normal:.1f} MB") +print(f"Streaming extraction: {memory_streaming:.1f} MB") +print(f"Memory savings: {memory_normal - memory_streaming:.1f} MB") +``` + +## Best Practices + +### For Development + +```python +# Fast compression for frequent builds +create_archive("build-artifacts.tzst", ["build/"], compression_level=1) +``` + +### For Backups + +```python +# Balanced compression for regular backups +create_archive("daily-backup.tzst", ["data/"], compression_level=6) +``` + +### For Distribution + +```python +# Higher compression for software distribution +create_archive("software-package.tzst", ["app/"], compression_level=9) +``` + +### For Archival Storage + +```python +# Maximum compression for long-term storage +create_archive("archive-2024.tzst", ["historical-data/"], compression_level=22) +``` + +## Hardware Considerations + +### CPU Usage + +- Higher compression levels use more CPU but for shorter time periods +- Modern multi-core systems handle zstd compression very efficiently +- Consider system load when choosing compression levels + +### Memory Usage + +- Streaming mode: ~16-32 MB memory usage regardless of archive size +- Normal mode: Memory usage proportional to archive size +- Use streaming for archives >100 MB or on memory-constrained systems + +### Storage + +- SSDs benefit from higher compression (less I/O) +- HDDs may prefer lower compression levels (CPU vs I/O trade-off) +- Network storage benefits from higher compression (bandwidth savings) + +## Integration with Build Systems + +### Makefile Example + +```makefile +# Fast compression for development +build-dev: + tzst a build-dev.tzst build/ -l 1 + +# Production compression +build-prod: + tzst a build-prod.tzst build/ -l 9 + +# CI/CD artifacts +artifacts: + tzst a artifacts.tzst dist/ logs/ -l 6 +``` + +### GitHub Actions Example + +```yaml +- name: Create release archive + run: | + tzst a release-${{ github.ref_name }}.tzst \ + build/ docs/ \ + --compression-level 9 +``` + +This performance guide helps you choose the right settings for your specific use case and understand how tzst compares to alternative archive tools. diff --git a/docs/quickstart.md b/docs/quickstart.md index beef535..3d10b7d 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -19,7 +19,7 @@ This guide will get you up and running with tzst in just a few minutes. Choose your preferred installation method: -### Option 1: PyPI (Recommended) +### Option 1: PyPI ```bash pip install tzst @@ -31,16 +31,33 @@ Download the appropriate executable from [GitHub Releases](https://github.com/xi | Platform | Architecture | Download | |----------|--------------|----------| -| **🐧 Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | -| **🐧 Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | -| **🪟 Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | -| **🪟 Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | -| **🍎 macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | -| **🍎 macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | +| **Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | +| **Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | +| **Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | +| **Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | +| **macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | +| **macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | Extract the archive and add the executable to your PATH. -### Option 3: From Source +### Option 3: Using uvx (No Installation) + +Run tzst directly without installation using [uvx](https://docs.astral.sh/uv/): + +```bash +uvx tzst --help +uvx tzst a archive.tzst file1.txt file2.txt directory/ +uvx tzst x archive.tzst +``` + +This option is perfect for: + +- **One-time usage** - No permanent installation needed +- **Testing** - Try tzst without committing to installation +- **CI/CD pipelines** - Use tzst in automated workflows +- **Isolated environments** - Avoid dependency conflicts + +### Option 4: From Source ```bash git clone https://github.com/xixu-me/tzst.git @@ -54,6 +71,8 @@ pip install . ### Command Line Interface +> **Note**: Download the [standalone binary](installation) for the best performance and no Python dependency. Alternatively, use `uvx tzst` for running without installation. See [uv documentation](https://docs.astral.sh/uv/) for details. + The CLI provides four main operations: ```bash @@ -70,6 +89,25 @@ tzst l archive.tzst tzst t archive.tzst ``` +### Command Reference + +| Command | Aliases | Description | Streaming Support | +|---------|---------|-------------|-------------------| +| `a` | `add`, `create` | Create or add to archive | N/A | +| `x` | `extract` | Extract with full paths | `--streaming` | +| `e` | `extract-flat` | Extract without directory structure | `--streaming` | +| `l` | `list` | List archive contents | `--streaming` | +| `t` | `test` | Test archive integrity | `--streaming` | + +### CLI Options + +- `-v, --verbose`: Enable verbose output +- `-o, --output DIR`: Specify output directory (extract commands) +- `-l, --level LEVEL`: Set compression level 1-22 (create command) +- `--streaming`: Enable streaming mode for memory-efficient processing +- `--filter FILTER`: Security filter for extraction (data/tar/fully_trusted) +- `--no-atomic`: Disable atomic file operations (not recommended) + #### Create Archives ```bash @@ -183,6 +221,29 @@ extract_archive("untrusted.tzst", "safe-output/", filter="data") extract_archive("trusted.tzst", "output/", filter="tar") ``` +### Security Filters + +tzst provides three security filter options for extraction: + +```python +from tzst import extract_archive + +# Extract with maximum security (default) +extract_archive("archive.tzst", "output/", filter="data") + +# Extract with standard tar compatibility +extract_archive("archive.tzst", "output/", filter="tar") + +# Extract with full trust (dangerous - only for trusted archives) +extract_archive("archive.tzst", "output/", filter="fully_trusted") +``` + +**Security Filter Options:** + +- `data` (default): Most secure. Blocks dangerous files, absolute paths, and paths outside extraction directory +- `tar`: Standard tar compatibility. Blocks absolute paths and directory traversal +- `fully_trusted`: No security restrictions. Only use with completely trusted archives + ### Conflict Resolution ```python @@ -211,6 +272,56 @@ create_archive("best.tzst", files, compression_level=22) # Best compression extract_archive("huge-archive.tzst", "output/", streaming=True) ``` +### Streaming Mode + +For large archives (>100MB), use streaming mode to reduce memory usage: + +```python +# Memory-efficient operations +with TzstArchive("large-archive.tzst", "r", streaming=True) as archive: + contents = archive.list() + archive.extractall("output/") + is_valid = archive.test() +``` + +**Note**: Streaming mode has limitations - you cannot extract specific files or use random access operations. + +### File Extensions + +The library automatically handles file extensions with intelligent normalization: + +- `.tzst` - Primary extension for tar+zstandard archives +- `.tar.zst` - Alternative standard extension +- Auto-detection when opening existing archives +- Automatic extension addition when creating archives + +```python +from tzst import create_archive + +# These all create valid archives +create_archive("backup.tzst", files) # Creates backup.tzst +create_archive("backup.tar.zst", files) # Creates backup.tar.zst +create_archive("backup", files) # Creates backup.tzst +create_archive("backup.txt", files) # Creates backup.tzst (normalized) +``` + +### Atomic Operations + +All file creation operations use atomic file operations by default: + +- Archives created in temporary files first, then atomically moved +- Automatic cleanup if process is interrupted +- No risk of corrupted or incomplete archives +- Cross-platform compatibility + +```python +# Atomic operations enabled by default +create_archive("important.tzst", files) # Safe from interruption + +# Can be disabled if needed (not recommended) +create_archive("test.tzst", files, use_temp_file=False) +``` + ## Error Handling ```python @@ -250,80 +361,6 @@ with TzstArchive("data.tzst", "r") as archive: members = archive.getmembers() ``` -## Important Concepts - -### Compression Levels - -tzst supports compression levels from 1 to 22: - -- **Level 1-3**: Fast compression, larger files (good for temporary archives) -- **Level 4-6**: Balanced compression and speed (recommended for most use cases) -- **Level 7-15**: Higher compression, slower (good for long-term storage) -- **Level 16-22**: Maximum compression, much slower (for size-critical applications) - -```python -# Fast compression -create_archive("temp.tzst", files, compression_level=1) - -# Balanced (default) -create_archive("backup.tzst", files, compression_level=3) - -# High compression -create_archive("archive.tzst", files, compression_level=9) - -# Maximum compression -create_archive("minimal.tzst", files, compression_level=22) -``` - -### Security Filters - -tzst provides extraction filters to protect against malicious archives: - -```python -# Safe data extraction (default, recommended) -extract_archive("archive.tzst", "output/", filter="data") - -# Preserve more tar features but still secure -extract_archive("archive.tzst", "output/", filter="tar") - -# Full trust mode (use only with trusted archives) -extract_archive("archive.tzst", "output/", filter="fully_trusted") -``` - -### Streaming Mode - -For large archives (>100MB), use streaming mode to reduce memory usage: - -```python -# Memory-efficient operations -with TzstArchive("large-archive.tzst", "r", streaming=True) as archive: - contents = archive.list() - archive.extractall("output/") - is_valid = archive.test() -``` - -**Note**: Streaming mode has limitations - you cannot extract specific files or use random access operations. - -### Handling File Conflicts - -Handle file conflicts during extraction: - -```python -from tzst import ConflictResolution - -# Skip existing files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.SKIP) - -# Replace all existing files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.REPLACE_ALL) - -# Auto-rename conflicting files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.AUTO_RENAME_ALL) -``` - ## Common Patterns ### Backup Script @@ -359,17 +396,16 @@ def verify_archive(archive_path): # Test integrity if not test_archive(archive_path): - print("❌ Archive is corrupted!") + print("Archive is corrupted!") return False # List contents contents = list_archive(archive_path, verbose=True) total_size = sum(item['size'] for item in contents if item['is_file']) file_count = sum(1 for item in contents if item['is_file']) - - print(f"✅ Archive is valid") - print(f"📁 Files: {file_count}") - print(f"📦 Total size: {total_size / 1024 / 1024:.1f} MB") + print(f"Archive is valid") + print(f"Files: {file_count}") + print(f"Total size: {total_size / 1024 / 1024:.1f} MB") return True ```