diff --git a/.gitignore b/.gitignore index 6e88cb6..9dcbc3d 100644 --- a/.gitignore +++ b/.gitignore @@ -1,6 +1,7 @@ # project *.session *.csv +output/ plans/ # python diff --git a/README.md b/README.md index 1d287c4..be6c73a 100644 --- a/README.md +++ b/README.md @@ -69,6 +69,34 @@ configuration: python compare.py ``` +## Export a crawl to CSV + +Export one group's saved crawl from Redis: + +```bash +python export.py [group_id] [timecrawl] +``` + +When `timecrawl` is omitted, the latest crawl for the group is exported. When +both arguments are omitted, the first group in the stored configuration and its +latest crawl are used. An explicit crawl time must use `yyyymmddhhmmss` format: + +```bash +python export.py -1001234567890 20260724120000 +``` + +CSV files are written to: + +```text +output/-.csv +``` + +Each file contains the columns `id`, `username`, `first_name`, and `last_name`. +Formula-like Telegram text is prefixed with an apostrophe for safe spreadsheet +opening. On POSIX systems, the `output/` directory is created with `0700` +permissions and each CSV file is written with `0600` permissions. The `output/` +directory is ignored by Git. + ## Configuration | Variable | Description | diff --git a/common.py b/common.py index 307d14e..3acf5c1 100644 --- a/common.py +++ b/common.py @@ -73,6 +73,20 @@ def list_group_exports(group_id): ] +def latest_group_export_time(group_id): + """Return the latest Redis-key timestamp for one group, or None.""" + latest_time = None + for export_key in redis_client.scan_iter( + match=key('group', str(group_id), '*'), + ): + if not _is_group_export_key(export_key): + continue + run_time = export_key.rsplit(':', 1)[-1] + if latest_time is None or run_time > latest_time: + latest_time = run_time + return latest_time + + def get_group_export(group_id, run_time): """Return a group export at one run time, or None when missing.""" raw = redis_client.get(key('group', str(group_id), run_time)) diff --git a/docs/journals/260724-1025-redis-csv-export.md b/docs/journals/260724-1025-redis-csv-export.md new file mode 100644 index 0000000..e6a4d0f --- /dev/null +++ b/docs/journals/260724-1025-redis-csv-export.md @@ -0,0 +1,40 @@ +# Redis-to-CSV Export + +--- +date: 2026-07-24 10:25 +session: redis-csv-export +--- + +## Context + +Saved Telegram crawls lived only in Redis. The new command makes one crawl +available as a local CSV without contacting Telegram. + +## What Happened + +- Added `export.py [group_id] [timecrawl]`. +- Used the latest Redis-key timestamp when `timecrawl` is omitted. +- Used the first configured group when both arguments are omitted. +- Wrote `id`, `username`, `first_name`, and `last_name` to + `output/-.csv`. +- Added `output/` to `.gitignore` and documented the command. + +## Decisions + +- Validate CLI timestamps, Redis record metadata, and every member before + creating output. +- Fail on a corrupt newest crawl instead of silently exporting stale history. +- Neutralize spreadsheet-formula prefixes in Telegram text. +- Write through a temporary file and atomically replace the destination. +- Use `0700` directory and `0600` file permissions on POSIX systems. + +## Verification + +- All 18 unit tests passed. +- Python compilation and whitespace checks passed. +- CLI help works without `REDIS_URL`. +- Independent testing and review found no remaining issues. + +## Next Steps + +- Keep CSV validation and README documentation aligned if export fields change. diff --git a/export.py b/export.py new file mode 100644 index 0000000..5420ab2 --- /dev/null +++ b/export.py @@ -0,0 +1,169 @@ +import argparse +import csv +import os +import sys +import tempfile +from datetime import datetime +from pathlib import Path + +OUTPUT_DIR = Path('output') +CSV_FIELDS = ['id', 'username', 'first_name', 'last_name'] +FORMULA_PREFIXES = ('=', '+', '-', '@', '\t', '\r') + + +def parse_args(argv=None): + parser = argparse.ArgumentParser( + description='Export a Redis-backed Telegram group crawl to CSV.', + ) + parser.add_argument( + 'group_id', + type=int, + nargs='?', + help='Telegram group ID (defaults to the first configured group)', + ) + parser.add_argument( + 'timecrawl', + nargs='?', + type=parse_run_time, + help='Crawl time in yyyymmddhhmmss format (defaults to latest)', + ) + return parser.parse_args(argv) + + +def parse_run_time(value): + if not is_valid_run_time(value): + if len(value) != 14 or not value.isdigit(): + raise argparse.ArgumentTypeError( + 'timecrawl must use yyyymmddhhmmss format', + ) + raise argparse.ArgumentTypeError( + 'timecrawl must be a valid date and time', + ) + return value + + +def is_valid_run_time(value): + if not isinstance(value, str) or len(value) != 14 or not value.isdigit(): + return False + try: + datetime.strptime(value, '%Y%m%d%H%M%S') + except ValueError: + return False + return True + + +def main(argv=None): + args = parse_args(argv) + + group_id = args.group_id + try: + # Lazy imports keep `python export.py --help` working without REDIS_URL. + from common import get_group_export, latest_group_export_time + + if group_id is None: + from config import load_app_config + + group_id = load_app_config()['group_ids'][0] + + run_time = args.timecrawl + if run_time is None: + run_time = latest_group_export_time(group_id) + if run_time is None: + print(f'no exports found for group {group_id}', file=sys.stderr) + return 1 + record = get_group_export(group_id, run_time) + except Exception as exc: + print(f'could not read export configuration or history: {exc}', file=sys.stderr) + return 1 + + if record is None: + print( + f'export not found or invalid for group {group_id} at {run_time}', + file=sys.stderr, + ) + return 1 + + try: + members = validate_record(record, group_id, run_time) + output_path = write_csv(group_id, run_time, members) + except (OSError, ValueError, csv.Error) as exc: + print(f'could not export group {group_id} at {run_time}: {exc}', file=sys.stderr) + return 1 + + print(f'exported {len(members)} members to {output_path}') + return 0 + + +def validate_record(record, expected_group_id, expected_time): + if not is_valid_run_time(expected_time): + raise ValueError('Redis export key has an invalid crawl time') + if not isinstance(record, dict): + raise ValueError('Redis export is not an object') + if record.get('group_id') != expected_group_id: + raise ValueError('Redis export group ID does not match its key') + if record.get('time') != expected_time: + raise ValueError('Redis export time does not match its key') + + members = record.get('members') + if not isinstance(members, list): + raise ValueError('Redis export members must be a list') + + validated = [] + for index, member in enumerate(members): + if not isinstance(member, dict): + raise ValueError(f'member {index} is not an object') + member_id = member.get('id') + if not isinstance(member_id, int) or isinstance(member_id, bool): + raise ValueError(f'member {index} has an invalid ID') + + row = {'id': member_id} + for field in CSV_FIELDS[1:]: + value = member.get(field) + if value is not None and not isinstance(value, str): + raise ValueError(f'member {index} has an invalid {field}') + row[field] = escape_formula(value) + validated.append(row) + return validated + + +def escape_formula(value): + if value and ( + value[0] in ('\t', '\r', '\n') + or value.lstrip().startswith(FORMULA_PREFIXES) + ): + return f"'{value}" + return value + + +def write_csv(group_id, run_time, members, output_dir=None): + output_dir = output_dir or OUTPUT_DIR + output_path = output_dir / f'{group_id}-{run_time}.csv' + output_dir.mkdir(mode=0o700, parents=True, exist_ok=True) + output_dir.chmod(0o700) + + temp_path = None + try: + with tempfile.NamedTemporaryFile( + mode='w', + encoding='utf-8', + newline='', + dir=output_dir, + prefix=f'.{output_path.name}.', + delete=False, + ) as csv_file: + temp_path = Path(csv_file.name) + writer = csv.DictWriter(csv_file, fieldnames=CSV_FIELDS) + writer.writeheader() + writer.writerows(members) + os.chmod(temp_path, 0o600) + os.replace(temp_path, output_path) + except Exception: + if temp_path is not None: + temp_path.unlink(missing_ok=True) + raise + + return output_path + + +if __name__ == '__main__': + raise SystemExit(main()) diff --git a/tests/test_common.py b/tests/test_common.py new file mode 100644 index 0000000..a6c619b --- /dev/null +++ b/tests/test_common.py @@ -0,0 +1,40 @@ +import os +import unittest +from unittest.mock import MagicMock, patch + +os.environ.setdefault('REDIS_URL', 'redis://localhost') + +import common + + +class LatestGroupExportTimeTest(unittest.TestCase): + def test_uses_latest_timestamp_from_redis_keys_without_fetching_records(self): + redis_client = MagicMock() + redis_client.scan_iter.return_value = [ + 'telegram-export:group:-100123:20260724120000', + 'telegram-export:group:-100123:20260723120000', + ] + + with patch.object(common, 'redis_client', redis_client): + run_time = common.latest_group_export_time(-100123) + + self.assertEqual(run_time, '20260724120000') + redis_client.scan_iter.assert_called_once_with( + match='telegram-export:group:-100123:*', + ) + redis_client.get.assert_not_called() + + def test_ignores_malformed_keys(self): + redis_client = MagicMock() + redis_client.scan_iter.return_value = [ + 'telegram-export:group:-100123:not-a-time', + ] + + with patch.object(common, 'redis_client', redis_client): + run_time = common.latest_group_export_time(-100123) + + self.assertIsNone(run_time) + + +if __name__ == '__main__': + unittest.main() diff --git a/tests/test_export.py b/tests/test_export.py new file mode 100644 index 0000000..6657230 --- /dev/null +++ b/tests/test_export.py @@ -0,0 +1,337 @@ +import csv +import io +import types +import unittest +from contextlib import redirect_stderr, redirect_stdout +from pathlib import Path +from tempfile import TemporaryDirectory +from unittest.mock import MagicMock, patch + +import export as export_command + + +def export_record(group_id=-100123, run_time='20260724120000'): + return { + 'group_id': group_id, + 'title': 'Test Group', + 'time': run_time, + 'members': [ + { + 'id': 1, + 'username': 'alice', + 'first_name': 'Alice', + 'last_name': 'Example', + }, + { + 'id': 2, + 'username': None, + 'first_name': 'Bob', + 'last_name': None, + }, + ], + } + + +class ExportCommandTest(unittest.TestCase): + def test_exports_exact_group_and_time(self): + record = export_record() + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock(), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + stdout = io.StringIO() + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stdout(stdout), + ): + result = export_command.main(['-100123', '20260724120000']) + + output_path = output_dir / '-100123-20260724120000.csv' + self.assertEqual(result, 0) + common.get_group_export.assert_called_once_with( + -100123, + '20260724120000', + ) + common.latest_group_export_time.assert_not_called() + self.assertTrue(output_path.is_file()) + self.assertIn(str(output_path), stdout.getvalue()) + + with output_path.open(encoding='utf-8', newline='') as csv_file: + rows = list(csv.DictReader(csv_file)) + + self.assertEqual( + csv.DictReader(io.StringIO(output_path.read_text())).fieldnames, + export_command.CSV_FIELDS, + ) + self.assertEqual(rows[0], { + 'id': '1', + 'username': 'alice', + 'first_name': 'Alice', + 'last_name': 'Example', + }) + self.assertEqual(rows[1]['username'], '') + self.assertEqual(rows[1]['last_name'], '') + + def test_uses_latest_export_when_time_is_omitted(self): + latest = export_record(run_time='20260724120000') + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=latest), + latest_group_export_time=MagicMock( + return_value='20260724120000', + ), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stdout(io.StringIO()), + ): + result = export_command.main(['-100123']) + + self.assertEqual(result, 0) + common.latest_group_export_time.assert_called_once_with(-100123) + common.get_group_export.assert_called_once_with( + -100123, + '20260724120000', + ) + self.assertTrue( + (output_dir / '-100123-20260724120000.csv').is_file(), + ) + + def test_uses_first_configured_group_when_both_args_are_omitted(self): + record = export_record(group_id=-100999) + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock( + return_value='20260724120000', + ), + ) + config = types.SimpleNamespace( + load_app_config=MagicMock(return_value={'group_ids': [-100999]}), + ) + + with TemporaryDirectory() as temp_dir: + with ( + patch.dict('sys.modules', {'common': common, 'config': config}), + patch.object( + export_command, + 'OUTPUT_DIR', + Path(temp_dir) / 'output', + ), + redirect_stdout(io.StringIO()), + ): + result = export_command.main([]) + + self.assertEqual(result, 0) + config.load_app_config.assert_called_once_with() + common.latest_group_export_time.assert_called_once_with(-100999) + common.get_group_export.assert_called_once_with( + -100999, + '20260724120000', + ) + + def test_default_group_config_failure_returns_clean_error(self): + common = types.SimpleNamespace( + get_group_export=MagicMock(), + latest_group_export_time=MagicMock(), + ) + config = types.SimpleNamespace( + load_app_config=MagicMock( + side_effect=RuntimeError('Redis unavailable'), + ), + ) + stderr = io.StringIO() + + with ( + patch.dict('sys.modules', {'common': common, 'config': config}), + redirect_stderr(stderr), + ): + result = export_command.main([]) + + self.assertEqual(result, 1) + self.assertIn('could not read export configuration', stderr.getvalue()) + common.latest_group_export_time.assert_not_called() + common.get_group_export.assert_not_called() + + def test_common_import_failure_returns_clean_error(self): + stderr = io.StringIO() + real_import = __import__ + + def fail_common_import(name, *args, **kwargs): + if name == 'common': + raise ValueError('Redis URL is invalid') + return real_import(name, *args, **kwargs) + + with ( + patch('builtins.__import__', side_effect=fail_common_import), + redirect_stderr(stderr), + ): + result = export_command.main(['-100123', '20260724120000']) + + self.assertEqual(result, 1) + self.assertIn('could not read export configuration', stderr.getvalue()) + + def test_missing_exact_export_returns_error_without_creating_output(self): + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=None), + latest_group_export_time=MagicMock(), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + stderr = io.StringIO() + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stderr(stderr), + ): + result = export_command.main(['-100123', '20260724120000']) + + self.assertEqual(result, 1) + self.assertFalse(output_dir.exists()) + self.assertIn('export not found', stderr.getvalue()) + + def test_missing_group_history_returns_error_without_creating_output(self): + common = types.SimpleNamespace( + get_group_export=MagicMock(), + latest_group_export_time=MagicMock(return_value=None), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + stderr = io.StringIO() + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stderr(stderr), + ): + result = export_command.main(['-100123']) + + self.assertEqual(result, 1) + self.assertFalse(output_dir.exists()) + self.assertIn('no exports found', stderr.getvalue()) + + def test_rejects_invalid_time_before_loading_redis(self): + for invalid_time in ( + '../unsafe', + '20261301120000', + '20260230010101', + '20260724246000', + ): + with self.subTest(invalid_time=invalid_time): + with ( + patch.dict('sys.modules', {'common': None}), + self.assertRaises(SystemExit), + ): + export_command.main(['-100123', invalid_time]) + + def test_rejects_mismatched_record_without_creating_output(self): + record = export_record(group_id='../escaped') + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock(), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + stderr = io.StringIO() + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stderr(stderr), + ): + result = export_command.main(['-100123', '20260724120000']) + + self.assertEqual(result, 1) + self.assertFalse(output_dir.exists()) + self.assertIn('group ID does not match', stderr.getvalue()) + + def test_rejects_impossible_time_from_redis_key(self): + record = export_record(run_time='99999999999999') + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock( + return_value='99999999999999', + ), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + stderr = io.StringIO() + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stderr(stderr), + ): + result = export_command.main(['-100123']) + + self.assertEqual(result, 1) + self.assertFalse(output_dir.exists()) + self.assertIn('invalid crawl time', stderr.getvalue()) + + def test_rejects_malformed_members_without_replacing_existing_csv(self): + record = export_record() + record['members'] = None + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock(), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + output_dir.mkdir() + output_path = output_dir / '-100123-20260724120000.csv' + output_path.write_text('existing export', encoding='utf-8') + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stderr(io.StringIO()), + ): + result = export_command.main(['-100123', '20260724120000']) + + self.assertEqual(result, 1) + self.assertEqual( + output_path.read_text(encoding='utf-8'), + 'existing export', + ) + + def test_escapes_spreadsheet_formulas_and_uses_private_permissions(self): + record = export_record() + record['members'][0]['first_name'] = ' =HYPERLINK("https://example.com")' + record['members'][0]['last_name'] = '\n+SUM(1,1)' + common = types.SimpleNamespace( + get_group_export=MagicMock(return_value=record), + latest_group_export_time=MagicMock(), + ) + + with TemporaryDirectory() as temp_dir: + output_dir = Path(temp_dir) / 'output' + with ( + patch.dict('sys.modules', {'common': common}), + patch.object(export_command, 'OUTPUT_DIR', output_dir), + redirect_stdout(io.StringIO()), + ): + result = export_command.main(['-100123', '20260724120000']) + + output_path = output_dir / '-100123-20260724120000.csv' + with output_path.open(encoding='utf-8', newline='') as csv_file: + rows = list(csv.DictReader(csv_file)) + + self.assertEqual(result, 0) + self.assertEqual( + rows[0]['first_name'], + '\' =HYPERLINK("https://example.com")', + ) + self.assertEqual(rows[0]['last_name'], "'\n+SUM(1,1)") + self.assertEqual(output_dir.stat().st_mode & 0o777, 0o700) + self.assertEqual(output_path.stat().st_mode & 0o777, 0o600) + + +if __name__ == '__main__': + unittest.main()