From 44843e003544a4e5905e79ae43a0b064c7c8edb3 Mon Sep 17 00:00:00 2001 From: Xi Xu Date: Fri, 6 Jun 2025 22:00:11 +0800 Subject: [PATCH] Update documentation and add new guides This commit updates multiple documentation files to improve clarity, remove emojis, and add new sections. Key changes include the addition of 'development.md' and 'performance.md', updates to the quickstart guide, and enhancements to the API and CLI documentation. These changes aim to provide better guidance for users and contributors. --- docs/404.md | 13 +- docs/api/cli.md | 12 +- docs/api/index.md | 8 +- docs/development.md | 353 ++++++++++++++++++++++++++++++++++++++++++++ docs/examples.md | 48 +++--- docs/index.md | 131 ++++++++++++++-- docs/performance.md | 278 ++++++++++++++++++++++++++++++++++ docs/quickstart.md | 210 +++++++++++++++----------- 8 files changed, 916 insertions(+), 137 deletions(-) create mode 100644 docs/development.md create mode 100644 docs/performance.md diff --git a/docs/404.md b/docs/404.md index afcbf70..261c6f3 100644 --- a/docs/404.md +++ b/docs/404.md @@ -11,7 +11,7 @@ myst: # 404 - Page Not Found -## 🔍 Oops! The page you're looking for doesn't exist +## Oops! The page you're looking for doesn't exist The URL you requested could not be found in the tzst documentation. This might happen if: @@ -20,14 +20,17 @@ The URL you requested could not be found in the tzst documentation. This might h - There's a typo in the URL - The page has been removed -## 🏠 Where would you like to go? +## Where would you like to go? ### Popular Pages - **{doc}`index`** - Documentation homepage - **{doc}`quickstart`** - Get started with tzst +- **{doc}`performance`** - Learn about performance optimizations - **{doc}`examples`** - Practical examples and use cases - **{doc}`api/index`** - Complete API reference +- **{doc}`development`** - Contributing to tzst +- **{ref}`genindex`** - Index of all documented items ### Quick Navigation @@ -36,7 +39,7 @@ The URL you requested could not be found in the tzst documentation. This might h - **Security Features** - {ref}`security-and-filtering` - **Error Handling** - {ref}`error-handling` -## 🔎 Search Documentation +## Search Documentation Use the search box in the top navigation to find what you're looking for, or browse through these sections: @@ -53,13 +56,13 @@ Use the search box in the top navigation to find what you're looking for, or bro - **Integration Examples** - {ref}`integration-examples` - **Performance Optimization** - {ref}`performance-optimization` -## 📚 Additional Resources +## Additional Resources - [GitHub Repository](https://github.com/xixu-me/tzst) - Source code and issue tracker - [PyPI Package](https://pypi.org/project/tzst/) - Download and installation - [Release Notes](https://github.com/xixu-me/tzst/releases) - Latest updates and changes -## 🐛 Report an Issue +## Report an Issue If you believe this is a broken link within our documentation, please [report it on GitHub](https://github.com/xixu-me/tzst/issues). diff --git a/docs/api/cli.md b/docs/api/cli.md index 7f01546..e07dcc0 100644 --- a/docs/api/cli.md +++ b/docs/api/cli.md @@ -25,17 +25,17 @@ The command-line interface module provides comprehensive functionality for the t The tzst CLI provides a powerful command-line interface for archive operations with intuitive commands and comprehensive options. The interface is designed for both interactive use and scripting, with robust error handling and user-friendly output. -### 🔧 Core Commands +### Core Commands | Command | Aliases | Description | Streaming Support | |---------|---------|-------------|-------------------| | `a` | `add`, `create` | Create or add to archive | N/A | -| `x` | `extract` | Extract with full paths | ✓ `--streaming` | -| `e` | `extract-flat` | Extract without directory structure | ✓ `--streaming` | -| `l` | `list` | List archive contents | ✓ `--streaming` | -| `t` | `test` | Test archive integrity | ✓ `--streaming` | +| `x` | `extract` | Extract with full paths | `--streaming` | +| `e` | `extract-flat` | Extract without directory structure | `--streaming` | +| `l` | `list` | List archive contents | `--streaming` | +| `t` | `test` | Test archive integrity | `--streaming` | -### 🎯 Key Features +### Key Features - **Intuitive Commands**: Simple, memorable command aliases (a, x, e, l, t) - **Streaming Support**: Memory-efficient processing for large archives diff --git a/docs/api/index.md b/docs/api/index.md index a3c11b8..308a71a 100644 --- a/docs/api/index.md +++ b/docs/api/index.md @@ -104,26 +104,26 @@ Exception hierarchy for comprehensive error handling and debugging support. ## Key Features -### 🛡️ Security First +### Security First - Built-in path traversal protection - Multiple security filter options - Safe extraction by default -### ⚡ High Performance +### High Performance - Zstandard compression with configurable levels - Streaming support for large archives - Memory-efficient operations -### 🔧 Developer Friendly +### Developer Friendly - Clean, Pythonic API - Comprehensive error handling - Context manager support - Extensive documentation and examples -### 🌐 Cross-Platform +### Cross-Platform - Works on Windows, macOS, and Linux - Consistent behavior across platforms diff --git a/docs/development.md b/docs/development.md new file mode 100644 index 0000000..10f0556 --- /dev/null +++ b/docs/development.md @@ -0,0 +1,353 @@ +# Development Guide + +This guide provides comprehensive information for developers contributing to or working with the tzst library. + +## Setting up Development Environment + +This project uses modern Python packaging standards: + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install -e .[dev] +``` + +The development installation includes all necessary tools: + +- **pytest** - Testing framework +- **ruff** - Linting and formatting +- **coverage** - Code coverage analysis +- **sphinx** - Documentation generation + +## Running Tests + +### Basic Test Commands + +```bash +# Run all tests +python -m pytest + +# Run tests with coverage +pytest --cov=tzst --cov-report=html + +# Or use the simpler command (coverage settings are in pyproject.toml) +pytest + +# Run with verbose output +python -m pytest -v + +# Run specific test file +python -m pytest tests/test_core.py + +# Run integration tests only +python -m pytest -m integration +``` + +### Test Structure + +- **Unit tests**: Test individual functions and methods +- **Integration tests**: Test component interactions +- **CLI tests**: Test command-line interface +- **Platform-specific tests**: Test OS-specific functionality + +### Writing Tests + +1. **Use descriptive test names:** + + ```python + def test_create_archive_with_compression_level_9(): + ``` + +2. **Use fixtures for common test data:** + + ```python + def test_extract_archive(sample_archive_path, temp_dir): + ``` + +3. **Test edge cases:** + - Empty files + - Large files + - Invalid inputs + - Corrupted archives + +4. **Add markers for test categorization:** + + ```python + @pytest.mark.integration + def test_full_archive_workflow(): + ``` + +## Code Quality + +### Running Code Style Tools + +```bash +# Check code quality +ruff check src tests + +# Fix auto-fixable issues +ruff check --fix src tests + +# Format code +ruff format src tests + +# Check formatting without making changes +ruff format --check src tests +``` + +### Configuration + +Settings are defined in `pyproject.toml`: + +- Line length: 88 characters +- Target Python version: 3.12+ +- Import sorting with isort +- Quote style: double quotes + +### Code Style Guidelines + +1. **Follow PEP 8** with project-specific modifications +2. **Use type hints** for all public APIs +3. **Write docstrings** for classes and public methods +4. **Keep functions focused** and reasonably sized +5. **Use meaningful variable names** +6. **Add comments** for complex logic + +## Documentation + +### Building Documentation + +```bash +# Navigate to docs directory +cd docs + +# Install documentation dependencies +pip install -r requirements.txt + +# Build HTML documentation +make html + +# On Windows, use: +make.bat html + +# View built documentation +# Open docs/_build/html/index.html in your browser +``` + +### Documentation Structure + +``` +docs/ +├── index.md # Main documentation landing page +├── quickstart.md # Getting started guide +├── performance.md # Performance guide and comparisons +├── examples.md # Usage examples +├── development.md # This development guide +├── api/ # API reference documentation +│ ├── index.md +│ ├── core.md +│ ├── cli.md +│ └── exceptions.md +├── conf.py # Sphinx configuration +└── requirements.txt # Documentation dependencies +``` + +### Writing Documentation + +- Use **MyST Markdown** format +- Include **code examples** for new features +- Add **cross-references** using proper syntax +- Test all **code snippets** to ensure they work + +## Project Structure + +``` +tzst/ +├── src/tzst/ # Main package source code +│ ├── __init__.py # Package initialization and exports +│ ├── __main__.py # CLI entry point +│ ├── cli.py # Command-line interface +│ ├── core.py # Core archive functionality +│ └── exceptions.py # Custom exceptions +├── tests/ # Test suite +│ ├── conftest.py # Pytest configuration and fixtures +│ ├── test_core.py # Core functionality tests +│ ├── test_cli.py # CLI tests +│ └── test_*.py # Additional test modules +├── docs/ # Documentation source +├── .github/ # GitHub workflows and templates +├── pyproject.toml # Project configuration +├── README.md # Project documentation +├── LICENSE # BSD 3-Clause License +└── CONTRIBUTING.md # Contribution guidelines +``` + +## 🤝 Contributing Workflow + +### 1. Making Changes + +#### Types of Contributions + +- **Bug fixes**: Fix issues in existing functionality +- **Features**: Add new capabilities to the library +- **Documentation**: Improve or add documentation +- **Tests**: Add or improve test coverage +- **Performance**: Optimize existing code +- **Security**: Address security vulnerabilities + +#### Branch Naming + +Use descriptive branch names: + +- `feature/add-streaming-mode` +- `fix/handle-corrupted-archives` +- `docs/improve-api-documentation` +- `test/add-compression-tests` + +### 2. Commit Messages + +Follow conventional commit format: + +``` +type(scope): description + +[optional body] + +[optional footer] +``` + +**Types:** + +- `feat`: New feature +- `fix`: Bug fix +- `docs`: Documentation changes +- `test`: Adding or modifying tests +- `refactor`: Code refactoring +- `perf`: Performance improvements +- `chore`: Build process or auxiliary tool changes + +**Examples:** + +``` +feat(core): add streaming compression support + +fix(cli): handle invalid archive paths gracefully + +docs(readme): update installation instructions +``` + +### 3. Pull Request Process + +1. **Create a feature branch:** + + ```bash + git checkout -b feature/your-feature-name + ``` + +2. **Make your changes** following the guidelines above + +3. **Add tests** for new functionality + +4. **Update documentation** if needed + +5. **Run the test suite:** + + ```bash + python -m pytest + ruff check . + ruff format --check . + ``` + +6. **Commit your changes:** + + ```bash + git add . + git commit -m "feat: add your feature description" + ``` + +7. **Push to your fork:** + + ```bash + git push origin feature/your-feature-name + ``` + +8. **Create a pull request** using the provided template + +### 4. Pull Request Guidelines + +- **Fill out the PR template** completely +- **Link related issues** using keywords (fixes #123) +- **Keep PRs focused** - one feature/fix per PR +- **Ensure all CI checks pass** +- **Respond to review feedback** promptly + +## Development Tips + +### Performance Considerations + +- Use streaming for large files +- Consider memory usage patterns +- Profile code for bottlenecks +- Test with various file sizes + +### Security Considerations + +- Validate all user inputs +- Use secure defaults (e.g., 'data' filter) +- Handle malicious archives safely +- Be cautious with file paths + +### Compatibility + +- Support Python 3.12+ +- Test on multiple platforms (Windows, macOS, Linux) +- Consider different filesystem behaviors +- Maintain backwards compatibility when possible + +## Release Process + +Releases are handled by maintainers: + +1. Update version in `src/tzst/__init__.py` +2. Create a release tag +3. Automated CI/CD publishes to PyPI + +## Getting Help + +### Resources + +- **Issues**: [GitHub Issues](https://github.com/xixu-me/tzst/issues) +- **Discussions**: Use GitHub Discussions for questions +- **Documentation**: Check the README and code comments + +### Reporting Issues + +When reporting bugs: + +1. **Use the bug report template** +2. **Provide a minimal reproduction case** +3. **Include system information** (OS, Python version) +4. **Attach relevant files** if possible (archives, logs) + +### Suggesting Features + +When suggesting features: + +1. **Use the feature request template** +2. **Explain the use case** and motivation +3. **Consider backwards compatibility** +4. **Provide implementation ideas** if you have them + +## Code of Conduct + +This project follows the principles of respectful collaboration. Please be kind, constructive, and professional in all interactions. + +## Recognition + +Contributors are recognized in several ways: + +- Listed in release notes for significant contributions +- Mentioned in README acknowledgments +- GitHub contributor statistics + +Thank you for contributing to tzst! Your efforts help make this library better for everyone. diff --git a/docs/examples.md b/docs/examples.md index 1a47103..c1052d8 100644 --- a/docs/examples.md +++ b/docs/examples.md @@ -89,14 +89,19 @@ for item in contents: ## Command Line Usage +> **Note**: Download the [standalone binary](https://github.com/xixu-me/tzst/releases) for the best performance and no Python dependency. Alternatively, use `uvx tzst` for running without installation. See [uv documentation](https://docs.astral.sh/uv/) for details. + ### Archive Creation Commands ```bash # Basic archive creation tzst a backup.tzst documents/ photos/ +# Or with uvx (no installation needed) +uvx tzst a backup.tzst documents/ photos/ # Create with specific compression level tzst a backup.tzst documents/ photos/ -l 6 +uvx tzst a backup.tzst documents/ photos/ -l 6 # Create from multiple sources tzst a complete-backup.tzst /home/user/documents /home/user/photos /etc/config @@ -110,6 +115,8 @@ tzst a backup.tzst documents/ photos/ -v ```bash # Extract to current directory tzst x backup.tzst +# Or with uvx +uvx tzst x backup.tzst # Extract to specific directory tzst x backup.tzst --output /restore/ @@ -129,6 +136,8 @@ tzst e backup.tzst --output flat-restore/ ```bash # List archive contents tzst l backup.tzst +# Or with uvx +uvx tzst l backup.tzst # List with detailed information tzst l backup.tzst --verbose @@ -551,23 +560,23 @@ def validate_and_repair_archive(archive_path): # Test basic integrity try: if test_archive(archive_path): - print("✅ Archive integrity test passed") + print("Archive integrity test passed") return True except Exception as e: - print(f"❌ Integrity test failed: {e}") + print(f"Integrity test failed: {e}") # Try to list contents try: contents = list_archive(archive_path) - print(f"📁 Archive contains {len(contents)} items") + print(f"Archive contains {len(contents)} items") # Try streaming mode if regular mode fails contents_streaming = list_archive(archive_path, streaming=True) if len(contents_streaming) != len(contents): - print("⚠️ Different results between modes - possible corruption") + print("Different results between modes - possible corruption") except Exception as e: - print(f"❌ Cannot list contents: {e}") + print(f"Cannot list contents: {e}") return False # Try partial extraction @@ -583,13 +592,13 @@ def validate_and_repair_archive(archive_path): archive.extract(member.name, backup_dir) extracted_count += 1 except Exception as e: - print(f"⚠️ Failed to extract {member.name}: {e}") + print(f"Failed to extract {member.name}: {e}") - print(f"✅ Recovered {extracted_count} files to {backup_dir}") + print(f"Recovered {extracted_count} files to {backup_dir}") return True except Exception as e: - print(f"❌ Recovery failed: {e}") + print(f"Recovery failed: {e}") return False # Example usage @@ -648,15 +657,15 @@ class BackupManager: # Validate the backup if test_archive(backup_path): file_size = backup_path.stat().st_size / (1024 * 1024) - print(f"✅ Backup created and validated: {file_size:.1f} MB") + print(f"Backup created and validated: {file_size:.1f} MB") return backup_path else: - print("❌ Backup validation failed!") + print("Backup validation failed!") backup_path.unlink() # Remove invalid backup return None except Exception as e: - print(f"❌ Backup failed: {e}") + print(f"Backup failed: {e}") return None def cleanup_old_backups(self): @@ -792,13 +801,13 @@ def archive_logs_by_date(log_dir, archive_dir, days_old=7): archived_files.append(log_file) file_size = archive_path.stat().st_size / 1024 - print(f"✅ Created {archive_name} ({file_size:.1f} KB)") + print(f"Created {archive_name} ({file_size:.1f} KB)") else: - print(f"❌ Archive validation failed for {archive_name}") + print(f"Archive validation failed for {archive_name}") archive_path.unlink() except Exception as e: - print(f"❌ Failed to archive logs for {date_key}: {e}") + print(f"Failed to archive logs for {date_key}: {e}") print(f"Archived {len(archived_files)} log files") @@ -856,9 +865,8 @@ class DataMigrator: with open(checksum_file, "w") as f: f.write(f"{checksum} {package_path.name}\n") - package_size = package_path.stat().st_size / (1024 * 1024) - print(f"✅ Package created: {package_size:.1f} MB") - print(f"📋 Checksum: {checksum}") + package_size = package_path.stat().st_size / (1024 * 1024) print(f"Package created: {package_size:.1f} MB") + print(f"Checksum: {checksum}") return package_path, checksum_file @@ -877,13 +885,13 @@ class DataMigrator: if expected_checksum != actual_checksum: raise RuntimeError(f"Checksum mismatch! Expected: {expected_checksum}, Got: {actual_checksum}") - print("✅ Checksum verification passed") + print("Checksum verification passed") # Test archive integrity if not test_archive(package_path): raise RuntimeError("Archive integrity check failed") - print("✅ Archive integrity verified") + print("Archive integrity verified") # Extract with conflict resolution destination.mkdir(parents=True, exist_ok=True) @@ -893,7 +901,7 @@ class DataMigrator: conflict_resolution="replace_all" # Overwrite for migration ) - print(f"✅ Package extracted to: {destination}") + print(f"Package extracted to: {destination}") # Example usage if __name__ == "__main__": diff --git a/docs/index.md b/docs/index.md index 13070b4..23acc21 100644 --- a/docs/index.md +++ b/docs/index.md @@ -11,6 +11,13 @@ myst: # tzst Documentation +[![codecov](https://codecov.io/gh/xixu-me/tzst/graph/badge.svg?token=2AIN1559WU)](https://codecov.io/gh/xixu-me/tzst) +[![CodeQL](https://github.com/xixu-me/tzst/actions/workflows/github-code-scanning/codeql/badge.svg)](https://github.com/xixu-me/tzst/actions/workflows/github-code-scanning/codeql) +[![CI/CD](https://github.com/xixu-me/tzst/actions/workflows/ci.yml/badge.svg)](https://github.com/xixu-me/tzst/actions/workflows/ci.yml) +[![PyPI - Version](https://img.shields.io/pypi/v/tzst)](https://pypi.org/project/tzst/) +[![GitHub License](https://img.shields.io/github/license/xixu-me/tzst)](LICENSE) +[![Sponsor](https://img.shields.io/badge/Sponsor-violet)](https://xi-xu.me/#sponsorships) + Welcome to **tzst**, the next-generation Python library engineered for modern archive management, leveraging cutting-edge Zstandard compression to deliver superior performance, security, and reliability. ```{toctree} @@ -18,8 +25,11 @@ Welcome to **tzst**, the next-generation Python library engineered for modern ar :caption: Contents: quickstart -api/index +performance examples +api/index +development +genindex ``` ```{toctree} @@ -41,25 +51,25 @@ README ## Key Features -### 🗜️ Advanced Compression +### Advanced Compression - **Zstandard Compression**: Best-in-class compression algorithm with configurable levels (1-22) - **Multiple Extensions**: Support for both `.tzst` and `.tar.zst` file extensions - **Streaming Support**: Memory-efficient processing for large archives -### 🔒 Security First +### Security First - **Safe by Default**: Uses 'data' filter for secure extraction without dangerous path traversal - **Multiple Filter Options**: Choose from 'data', 'tar', or 'fully_trusted' filters based on your security needs - **Atomic Operations**: All file operations use temporary files with atomic moves to prevent corruption -### 💻 Dual Interfaces +### Dual Interfaces - **Command Line**: Intuitive CLI with comprehensive options for batch operations - **Python API**: Clean, object-oriented interface for programmatic use - **Convenience Functions**: High-level functions for common operations -### ⚡ High Performance +### High Performance - **Optimized I/O**: Efficient buffering and streaming for large files - **Conflict Resolution**: Intelligent handling of file conflicts during extraction @@ -91,6 +101,49 @@ with TzstArchive("data.tzst", "r") as archive: pip install tzst ``` +### From GitHub Releases + +Download platform-specific standalone executables from [GitHub Releases](https://github.com/xixu-me/tzst/releases) - no Python installation required! + +#### Supported Platforms + +| Platform | Architecture | File | +|----------|-------------|------| +| **Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | +| **Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | +| **Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | +| **Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | +| **macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | +| **macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | + +#### Installation Steps + +1. **Download** the appropriate archive for your platform from the [latest releases page](https://github.com/xixu-me/tzst/releases/latest) +2. **Extract** the archive to get the `tzst` executable (or `tzst.exe` on Windows) +3. **Move** the executable to a directory in your PATH: + - **Linux/macOS**: `sudo mv tzst /usr/local/bin/` + - **Windows**: Add the directory containing `tzst.exe` to your PATH environment variable +4. **Verify** installation: `tzst --help` + +#### Benefits of Binary Installation + +- **No Python required** - Standalone executable +- **Faster startup** - No Python interpreter overhead +- **Easy deployment** - Single file distribution +- **Consistent behavior** - Bundled dependencies + +### Using uvx (No Installation) + +Run tzst directly without installation using [uvx](https://docs.astral.sh/uv/): + +```bash +uvx tzst --help +uvx tzst a archive.tzst file1.txt file2.txt directory/ +uvx tzst x archive.tzst +``` + +Perfect for one-time usage, testing, CI/CD pipelines, and isolated environments. + ### From Source ```bash @@ -99,19 +152,16 @@ cd tzst pip install . ``` -### Standalone Binaries - -Download platform-specific standalone executables from [GitHub Releases](https://github.com/xixu-me/tzst/releases) - no Python installation required! - ## Getting Started For a quick introduction, see the {doc}`quickstart` guide. For comprehensive usage examples, explore the {doc}`examples` section. ### Installation Options -1. **PyPI Installation** (Recommended): `pip install tzst` +1. **PyPI Installation**: `pip install tzst` 2. **Standalone Binaries**: Download from [GitHub Releases](https://github.com/xixu-me/tzst/releases) -3. **From Source**: Clone and install from repository +3. **uvx (No Installation)**: Run directly with `uvx tzst` +4. **From Source**: Clone and install from repository ### API Documentation @@ -121,12 +171,63 @@ Complete API documentation is available in the {doc}`api/index` section, coverin - {doc}`api/cli`: Command-line interface - {doc}`api/exceptions`: Error handling -## Indices and tables +## Development + +For comprehensive development information, see the {doc}`development` guide, which covers: + +- Setting up development environment +- Running tests and code quality checks +- Documentation building +- Contributing workflow and guidelines +- Project structure and best practices + +### Quick Start + +```bash +git clone https://github.com/xixu-me/tzst.git +cd tzst +pip install -e .[dev] +pytest +``` + +## Contributing + +We welcome contributions! Please read our [Contributing Guide](https://github.com/xixu-me/tzst/blob/main/CONTRIBUTING.md) for: + +- Development setup and project structure +- Code style guidelines and best practices +- Testing requirements and writing tests +- Pull request process and review workflow + +### Types of Contributions Welcome + +- **Bug fixes** - Fix issues in existing functionality +- **Features** - Add new capabilities to the library +- **Documentation** - Improve or add documentation +- **Tests** - Add or improve test coverage +- **Performance** - Optimize existing code +- **Security** - Address security vulnerabilities + +## Acknowledgments + +- [Meta Zstandard](https://github.com/facebook/zstd) for the excellent compression algorithm +- [python-zstandard](https://github.com/indygreg/python-zstandard) for Python bindings +- The Python community for inspiration and feedback + +## License + +Copyright © 2025 [Xi Xu](https://xi-xu.me). All rights reserved. + +Licensed under the [BSD 3-Clause](https://github.com/xixu-me/tzst/blob/main/LICENSE) license. + +## Documentation Guide 1. **{doc}`quickstart`** - Get up and running quickly with basic examples -2. **{doc}`examples`** - Comprehensive usage examples and patterns -3. **{doc}`api/index`** - Complete API reference documentation -4. **{ref}`genindex`** - Index of all documented items +2. **{doc}`performance`** - Performance optimization guide and comparisons +3. **{doc}`examples`** - Comprehensive usage examples and patterns +4. **{doc}`api/index`** - Complete API reference documentation +5. **{doc}`development`** - Development and contribution guidelines +6. **{ref}`genindex`** - Index of all documented items ## Requirements diff --git a/docs/performance.md b/docs/performance.md new file mode 100644 index 0000000..6677326 --- /dev/null +++ b/docs/performance.md @@ -0,0 +1,278 @@ +--- +myst: + html_meta: + description: "tzst Performance Guide - Compression level optimization, performance tips, and comparison with other archive tools" + keywords: "tzst performance, compression benchmarks, tar gzip comparison, archive performance optimization" + og:title: "tzst Performance Guide" + og:description: "Performance optimization tips and comparison with other archive tools for tzst" + twitter:title: "tzst Performance Guide" + twitter:description: "Performance optimization tips and comparison with other archive tools for tzst" +--- + +# Performance Guide + +This guide covers performance optimization techniques and provides detailed comparisons with other archive tools. + +## Performance Tips + +### 1. Compression Levels + +Choose the right compression level for your use case: + +- **Level 1-3**: Fast compression, larger files (good for temporary archives or real-time processing) +- **Level 3** (default): Optimal balance for most use cases +- **Level 6-9**: Higher compression, moderate speed (good for regular backups) +- **Level 15-22**: Maximum compression, slower (for long-term storage or bandwidth-limited scenarios) + +```python +from tzst import create_archive + +# For temporary files or frequent operations +create_archive("temp.tzst", files, compression_level=1) + +# Balanced default (recommended) +create_archive("backup.tzst", files, compression_level=3) + +# Long-term storage +create_archive("archive.tzst", files, compression_level=9) + +# Maximum compression for critical space savings +create_archive("minimal.tzst", files, compression_level=22) +``` + +### 2. Streaming + +Use streaming mode for archives larger than 100MB: + +```python +from tzst import extract_archive, list_archive, test_archive + +# Memory-efficient operations for large archives +extract_archive("large-backup.tzst", "restore/", streaming=True) +contents = list_archive("large-backup.tzst", streaming=True) +is_valid = test_archive("large-backup.tzst", streaming=True) +``` + +**Streaming Benefits:** + +- Significantly reduced memory usage +- Better performance for large archives +- Handles archives that don't fit in memory + +### 3. Batch Operations + +Add multiple files in a single session when possible: + +```python +from tzst import TzstArchive + +# Efficient: Single archive session +with TzstArchive("backup.tzst", "w") as archive: + archive.add("file1.txt") + archive.add("file2.txt") + archive.add("directory/", recursive=True) + +# Less efficient: Multiple separate operations +create_archive("backup1.tzst", ["file1.txt"]) +create_archive("backup2.tzst", ["file2.txt"]) +``` + +### 4. File Type Considerations + +- Already compressed files (`.jpg`, `.png`, `.mp4`, `.pdf`) won't compress much further +- Text files, source code, and logs compress very well +- Consider compression level based on your data types + +## Comparison with Other Tools + +### vs tar + gzip + +**tzst Advantages:** + +- **Better compression ratios**: 10-40% smaller archives +- **Faster decompression**: 2-3x faster extraction +- **Modern algorithm**: Better handling of various file types +- **Streaming support**: Better memory efficiency + +**When to use tar + gzip:** + +- Legacy system compatibility requirements +- Very old systems without zstd support + +### vs tar + xz + +**tzst Advantages:** + +- **Significantly faster compression**: 3-10x faster creation +- **Faster decompression**: 2-4x faster extraction +- **Better speed/compression trade-off**: Similar compression with much better speed +- **More compression levels**: Fine-grained control (22 levels vs 9) + +**When to use tar + xz:** + +- Maximum compression is critical and time is not a factor +- Systems that don't support zstd + +### vs zip + +**tzst Advantages:** + +- **Better compression**: 15-30% smaller archives +- **Preserves Unix permissions and metadata**: Full POSIX compatibility +- **Better streaming support**: Memory-efficient for large archives +- **Better directory handling**: Preserves directory structure and timestamps + +**When to use zip:** + +- Cross-platform compatibility with very old systems +- Individual file access without full extraction is required +- Windows-centric environments with no command-line tools + +## Benchmarking Examples + +### Compression Level Benchmark + +```python +import time +from pathlib import Path +from tzst import create_archive + +def benchmark_compression_levels(files, output_prefix="benchmark"): + """Compare different compression levels.""" + levels_to_test = [1, 3, 6, 9, 15, 22] + + results = [] + for level in levels_to_test: + output_file = f"{output_prefix}_level_{level}.tzst" + + # Measure compression time + start_time = time.time() + create_archive(output_file, files, compression_level=level) + compress_time = time.time() - start_time + + # Get file size + file_size = Path(output_file).stat().st_size + + results.append({ + 'level': level, + 'time': compress_time, + 'size': file_size, + 'size_mb': file_size / (1024 * 1024) + }) + + print(f"Level {level}: {compress_time:.2f}s, {file_size/1024/1024:.1f} MB") + + return results + +# Example usage +files = ["documents/", "projects/"] +results = benchmark_compression_levels(files) +``` + +### Memory Usage Comparison + +```python +import psutil +import os +from tzst import extract_archive + +def monitor_memory_usage(func, *args, **kwargs): + """Monitor memory usage during function execution.""" + process = psutil.Process(os.getpid()) + initial_memory = process.memory_info().rss / 1024 / 1024 # MB + + func(*args, **kwargs) + + peak_memory = process.memory_info().rss / 1024 / 1024 # MB + return peak_memory - initial_memory + +# Compare streaming vs non-streaming extraction +large_archive = "large-dataset.tzst" + +memory_normal = monitor_memory_usage(extract_archive, large_archive, "output1/") +memory_streaming = monitor_memory_usage(extract_archive, large_archive, "output2/", streaming=True) + +print(f"Normal extraction: {memory_normal:.1f} MB") +print(f"Streaming extraction: {memory_streaming:.1f} MB") +print(f"Memory savings: {memory_normal - memory_streaming:.1f} MB") +``` + +## Best Practices + +### For Development + +```python +# Fast compression for frequent builds +create_archive("build-artifacts.tzst", ["build/"], compression_level=1) +``` + +### For Backups + +```python +# Balanced compression for regular backups +create_archive("daily-backup.tzst", ["data/"], compression_level=6) +``` + +### For Distribution + +```python +# Higher compression for software distribution +create_archive("software-package.tzst", ["app/"], compression_level=9) +``` + +### For Archival Storage + +```python +# Maximum compression for long-term storage +create_archive("archive-2024.tzst", ["historical-data/"], compression_level=22) +``` + +## Hardware Considerations + +### CPU Usage + +- Higher compression levels use more CPU but for shorter time periods +- Modern multi-core systems handle zstd compression very efficiently +- Consider system load when choosing compression levels + +### Memory Usage + +- Streaming mode: ~16-32 MB memory usage regardless of archive size +- Normal mode: Memory usage proportional to archive size +- Use streaming for archives >100 MB or on memory-constrained systems + +### Storage + +- SSDs benefit from higher compression (less I/O) +- HDDs may prefer lower compression levels (CPU vs I/O trade-off) +- Network storage benefits from higher compression (bandwidth savings) + +## Integration with Build Systems + +### Makefile Example + +```makefile +# Fast compression for development +build-dev: + tzst a build-dev.tzst build/ -l 1 + +# Production compression +build-prod: + tzst a build-prod.tzst build/ -l 9 + +# CI/CD artifacts +artifacts: + tzst a artifacts.tzst dist/ logs/ -l 6 +``` + +### GitHub Actions Example + +```yaml +- name: Create release archive + run: | + tzst a release-${{ github.ref_name }}.tzst \ + build/ docs/ \ + --compression-level 9 +``` + +This performance guide helps you choose the right settings for your specific use case and understand how tzst compares to alternative archive tools. diff --git a/docs/quickstart.md b/docs/quickstart.md index beef535..3d10b7d 100644 --- a/docs/quickstart.md +++ b/docs/quickstart.md @@ -19,7 +19,7 @@ This guide will get you up and running with tzst in just a few minutes. Choose your preferred installation method: -### Option 1: PyPI (Recommended) +### Option 1: PyPI ```bash pip install tzst @@ -31,16 +31,33 @@ Download the appropriate executable from [GitHub Releases](https://github.com/xi | Platform | Architecture | Download | |----------|--------------|----------| -| **🐧 Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | -| **🐧 Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | -| **🪟 Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | -| **🪟 Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | -| **🍎 macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | -| **🍎 macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | +| **Linux** | x86_64 | `tzst-v{version}-linux-x86_64.zip` | +| **Linux** | ARM64 | `tzst-v{version}-linux-aarch64.zip` | +| **Windows** | x64 | `tzst-v{version}-windows-amd64.zip` | +| **Windows** | ARM64 | `tzst-v{version}-windows-arm64.zip` | +| **macOS** | Intel | `tzst-v{version}-macos-x86_64.zip` | +| **macOS** | Apple Silicon | `tzst-v{version}-macos-arm64.zip` | Extract the archive and add the executable to your PATH. -### Option 3: From Source +### Option 3: Using uvx (No Installation) + +Run tzst directly without installation using [uvx](https://docs.astral.sh/uv/): + +```bash +uvx tzst --help +uvx tzst a archive.tzst file1.txt file2.txt directory/ +uvx tzst x archive.tzst +``` + +This option is perfect for: + +- **One-time usage** - No permanent installation needed +- **Testing** - Try tzst without committing to installation +- **CI/CD pipelines** - Use tzst in automated workflows +- **Isolated environments** - Avoid dependency conflicts + +### Option 4: From Source ```bash git clone https://github.com/xixu-me/tzst.git @@ -54,6 +71,8 @@ pip install . ### Command Line Interface +> **Note**: Download the [standalone binary](installation) for the best performance and no Python dependency. Alternatively, use `uvx tzst` for running without installation. See [uv documentation](https://docs.astral.sh/uv/) for details. + The CLI provides four main operations: ```bash @@ -70,6 +89,25 @@ tzst l archive.tzst tzst t archive.tzst ``` +### Command Reference + +| Command | Aliases | Description | Streaming Support | +|---------|---------|-------------|-------------------| +| `a` | `add`, `create` | Create or add to archive | N/A | +| `x` | `extract` | Extract with full paths | `--streaming` | +| `e` | `extract-flat` | Extract without directory structure | `--streaming` | +| `l` | `list` | List archive contents | `--streaming` | +| `t` | `test` | Test archive integrity | `--streaming` | + +### CLI Options + +- `-v, --verbose`: Enable verbose output +- `-o, --output DIR`: Specify output directory (extract commands) +- `-l, --level LEVEL`: Set compression level 1-22 (create command) +- `--streaming`: Enable streaming mode for memory-efficient processing +- `--filter FILTER`: Security filter for extraction (data/tar/fully_trusted) +- `--no-atomic`: Disable atomic file operations (not recommended) + #### Create Archives ```bash @@ -183,6 +221,29 @@ extract_archive("untrusted.tzst", "safe-output/", filter="data") extract_archive("trusted.tzst", "output/", filter="tar") ``` +### Security Filters + +tzst provides three security filter options for extraction: + +```python +from tzst import extract_archive + +# Extract with maximum security (default) +extract_archive("archive.tzst", "output/", filter="data") + +# Extract with standard tar compatibility +extract_archive("archive.tzst", "output/", filter="tar") + +# Extract with full trust (dangerous - only for trusted archives) +extract_archive("archive.tzst", "output/", filter="fully_trusted") +``` + +**Security Filter Options:** + +- `data` (default): Most secure. Blocks dangerous files, absolute paths, and paths outside extraction directory +- `tar`: Standard tar compatibility. Blocks absolute paths and directory traversal +- `fully_trusted`: No security restrictions. Only use with completely trusted archives + ### Conflict Resolution ```python @@ -211,6 +272,56 @@ create_archive("best.tzst", files, compression_level=22) # Best compression extract_archive("huge-archive.tzst", "output/", streaming=True) ``` +### Streaming Mode + +For large archives (>100MB), use streaming mode to reduce memory usage: + +```python +# Memory-efficient operations +with TzstArchive("large-archive.tzst", "r", streaming=True) as archive: + contents = archive.list() + archive.extractall("output/") + is_valid = archive.test() +``` + +**Note**: Streaming mode has limitations - you cannot extract specific files or use random access operations. + +### File Extensions + +The library automatically handles file extensions with intelligent normalization: + +- `.tzst` - Primary extension for tar+zstandard archives +- `.tar.zst` - Alternative standard extension +- Auto-detection when opening existing archives +- Automatic extension addition when creating archives + +```python +from tzst import create_archive + +# These all create valid archives +create_archive("backup.tzst", files) # Creates backup.tzst +create_archive("backup.tar.zst", files) # Creates backup.tar.zst +create_archive("backup", files) # Creates backup.tzst +create_archive("backup.txt", files) # Creates backup.tzst (normalized) +``` + +### Atomic Operations + +All file creation operations use atomic file operations by default: + +- Archives created in temporary files first, then atomically moved +- Automatic cleanup if process is interrupted +- No risk of corrupted or incomplete archives +- Cross-platform compatibility + +```python +# Atomic operations enabled by default +create_archive("important.tzst", files) # Safe from interruption + +# Can be disabled if needed (not recommended) +create_archive("test.tzst", files, use_temp_file=False) +``` + ## Error Handling ```python @@ -250,80 +361,6 @@ with TzstArchive("data.tzst", "r") as archive: members = archive.getmembers() ``` -## Important Concepts - -### Compression Levels - -tzst supports compression levels from 1 to 22: - -- **Level 1-3**: Fast compression, larger files (good for temporary archives) -- **Level 4-6**: Balanced compression and speed (recommended for most use cases) -- **Level 7-15**: Higher compression, slower (good for long-term storage) -- **Level 16-22**: Maximum compression, much slower (for size-critical applications) - -```python -# Fast compression -create_archive("temp.tzst", files, compression_level=1) - -# Balanced (default) -create_archive("backup.tzst", files, compression_level=3) - -# High compression -create_archive("archive.tzst", files, compression_level=9) - -# Maximum compression -create_archive("minimal.tzst", files, compression_level=22) -``` - -### Security Filters - -tzst provides extraction filters to protect against malicious archives: - -```python -# Safe data extraction (default, recommended) -extract_archive("archive.tzst", "output/", filter="data") - -# Preserve more tar features but still secure -extract_archive("archive.tzst", "output/", filter="tar") - -# Full trust mode (use only with trusted archives) -extract_archive("archive.tzst", "output/", filter="fully_trusted") -``` - -### Streaming Mode - -For large archives (>100MB), use streaming mode to reduce memory usage: - -```python -# Memory-efficient operations -with TzstArchive("large-archive.tzst", "r", streaming=True) as archive: - contents = archive.list() - archive.extractall("output/") - is_valid = archive.test() -``` - -**Note**: Streaming mode has limitations - you cannot extract specific files or use random access operations. - -### Handling File Conflicts - -Handle file conflicts during extraction: - -```python -from tzst import ConflictResolution - -# Skip existing files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.SKIP) - -# Replace all existing files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.REPLACE_ALL) - -# Auto-rename conflicting files -extract_archive("archive.tzst", "output/", - conflict_resolution=ConflictResolution.AUTO_RENAME_ALL) -``` - ## Common Patterns ### Backup Script @@ -359,17 +396,16 @@ def verify_archive(archive_path): # Test integrity if not test_archive(archive_path): - print("❌ Archive is corrupted!") + print("Archive is corrupted!") return False # List contents contents = list_archive(archive_path, verbose=True) total_size = sum(item['size'] for item in contents if item['is_file']) file_count = sum(1 for item in contents if item['is_file']) - - print(f"✅ Archive is valid") - print(f"📁 Files: {file_count}") - print(f"📦 Total size: {total_size / 1024 / 1024:.1f} MB") + print(f"Archive is valid") + print(f"Files: {file_count}") + print(f"Total size: {total_size / 1024 / 1024:.1f} MB") return True ```