最新国产好看的视频,伊人天堂AV在线,国产Aaaaaa视频,蜜臀视频在线观看一区,人妻av色图,密臀久久久精品影片,青青视频免费观看毛片,久草在线观看视,国产三级精品色情在线

Python實(shí)現(xiàn)根據(jù)文件后綴統(tǒng)計(jì)文件大小并找出文件位置

 更新時(shí)間:2026年05月01日 07:34:41   作者:time_error  
這篇文章主要和大家分享了一個(gè)Python腳本,主要用于統(tǒng)計(jì)指定文件夾內(nèi)特定后綴文件的數(shù)量、大小及路徑,并支持多種格式導(dǎo)出,文中的示例代碼講解詳細(xì),感興趣的小伙伴可以了解下

需求分析

這個(gè)腳本支持掃描并統(tǒng)計(jì)某個(gè)文件夾內(nèi),以某個(gè)后綴結(jié)尾的,比如sql,java(以使用參數(shù)為準(zhǔn))。并且可以看到其中的文件大小和完整文件路徑,還支持以CSV、XLSX等格式導(dǎo)出等。這個(gè)腳本適用于統(tǒng)計(jì)本地代碼行數(shù)之后,需要查看某個(gè)編程語(yǔ)言,比如Java寫(xiě)了多少文件、每個(gè)文件大小是多少、文件路徑在哪里等操作。

使用教程

我們使用腳本時(shí),輸入想要檢測(cè)的路徑、文件后綴、以及是否導(dǎo)出文件、導(dǎo)出文件的命名(文件數(shù)量過(guò)多時(shí)使用)等就好。 比如python list_zip_contents.py -tp “C:\Users\your_name\Pictures\” -a -o rpng.csv ,演示時(shí)為best.csv。

最后,用編輯器打開(kāi)即可。

詳細(xì)代碼

# -*- coding: utf-8 -*-
import zipfile
import os
import argparse

EXTENSIONS = {
    'sql': '.sql', 'java': '.java', 'c': '.c', 'cpp': '.cpp', 'h': '.h',
    'py': '.py', 'js': '.js', 'html': '.html', 'css': '.css', 'markdown': '.md',
    'txt': '.txt', 'xml': '.xml', 'yaml': '.yaml', 'yml': '.yml', 'json': '.json',
    'properties': '.properties', 'go': '.go', 'rs': '.rs', 'rb': '.rb',
    'php': '.php', 'swift': '.swift', 'kt': '.kt', 'cs': '.cs', 'lua': '.lua',
    'sh': '.sh', 'bat': '.bat', 'ps1': '.ps1', 'sql': '.sql'
}

def split_archive_path(path):
    lower_path = path.lower()
    all_extensions = ['.tar.gz', '.tar.bz2', '.tar.zst', '.tgz', '.zip', '.7z', '.rar', '.tar', '.gz', '.bz2']
    candidates = []
    for ext in all_extensions:
        ext_lower = ext.lower()
        idx = lower_path.rfind(ext_lower)
        if idx >= 0:
            candidates.append((idx, len(ext), ext, path[:idx + len(ext)], path[idx + len(ext):].lstrip('\\').lstrip('/')))
    if candidates:
        candidates.sort(key=lambda x: (x[0], -x[1]))
        _, _, _, archive_path, sub_path = candidates[0]
        if sub_path:
            return archive_path, sub_path
        else:
            return archive_path, ''
    return path, ''

def is_archive(path):
    return path.lower().endswith(('.zip', '.7z', '.rar', '.tar', '.tar.gz', '.tar.bz2', '.tgz', '.gz', '.bz2'))

def list_files_in_path(target_path, filter_exts, show_all):
    if os.path.isdir(target_path):
        return list_files_in_dir(target_path, filter_exts, show_all)
    elif os.path.isfile(target_path) and is_archive(target_path):
        return list_files_in_archive(target_path, filter_exts, show_all)
    else:
        archive_path, sub_path = split_archive_path(target_path)
        if os.path.isfile(archive_path) and is_archive(archive_path) and sub_path:
            return list_files_in_archive_subdir(archive_path, sub_path, filter_exts, show_all)
        else:
            return {'error': f'Not a valid directory or zip file - {target_path}', 'files': [], 'header': ''}

def list_files_in_archive_subdir(zip_path, sub_path, filter_exts, show_all):
    import tempfile
    temp_dir = tempfile.mkdtemp()
    try:
        if zip_path.lower().endswith('.zip'):
            with zipfile.ZipFile(zip_path, 'r') as z:
                for member in z.namelist():
                    normalized = member.replace('/', os.sep).replace('\\', os.sep)
                    sub_path_normalized = sub_path.replace('/', os.sep).replace('\\', os.sep)
                    if normalized.startswith(sub_path_normalized + os.sep) or normalized == sub_path_normalized:
                        z.extract(member, temp_dir)
        elif zip_path.lower().endswith('.7z'):
            import py7zr
            with py7zr.SevenZipFile(zip_path, 'r') as sz:
                sz.extractall(temp_dir)
        else:
            return {'error': f'Unsupported archive format - {zip_path}', 'files': [], 'header': ''}
        extract_subdir = temp_dir
        for item in os.listdir(temp_dir):
            if sub_path.replace('/', os.sep).replace('\\', os.sep).endswith(item) or item.endswith('.zip') or item.endswith('.7z'):
                potential = os.path.join(temp_dir, item)
                if os.path.isdir(potential):
                    extract_subdir = potential
                    break
        return list_files_recursive(extract_subdir, filter_exts, show_all, zip_path, sub_path)
    finally:
        import shutil
        try:
            shutil.rmtree(temp_dir)
        except:
            pass

def list_files_recursive(dir_path, filter_exts, show_all, archive_path=None, sub_path=None):
    files_found = []
    dirs_to_scan = [dir_path]

    while dirs_to_scan:
        current_dir = dirs_to_scan.pop()
        for root, dirs, files in os.walk(current_dir):
            for filename in files:
                filepath = os.path.join(root, filename)
                try:
                    size = os.path.getsize(filepath)
                except:
                    size = 0
                rel_path = os.path.relpath(filepath, dir_path)

                ext = filename.lower()
                if ext.endswith(('.zip', '.7z', '.rar', '.tar', '.tar.gz', '.tgz', '.gz', '.bz2')):
                    nested_result = try_extract_nested(filepath, filter_exts, show_all)
                    if nested_result:
                        files_found.extend(nested_result['files'])
                    continue

                if show_all:
                    files_found.append((rel_path, size))
                else:
                    file_ext = os.path.splitext(filename)[1].lower()
                    if file_ext in filter_exts:
                        files_found.append((rel_path, size))

    if archive_path and sub_path:
        header = f'Archive: {archive_path} / {sub_path}'
    elif archive_path:
        header = f'Archive: {archive_path}'
    else:
        header = f'Directory: {dir_path}'
    header += f'\nTotal files: {len(files_found)}\n' + '=' * 80
    return {'files': sorted(files_found), 'header': header}

def try_extract_nested(archive_file, filter_exts, show_all):
    import tempfile
    temp_dir = tempfile.mkdtemp()
    try:
        ext = archive_file.lower()
        if ext.endswith('.zip'):
            with zipfile.ZipFile(archive_file, 'r') as z:
                z.extractall(temp_dir)
        elif ext.endswith('.7z'):
            import py7zr
            with py7zr.SevenZipFile(archive_file, 'r') as sz:
                sz.extractall(temp_dir)
        else:
            return None

        files_found = []
        for root, dirs, files in os.walk(temp_dir):
            for filename in files:
                filepath = os.path.join(root, filename)
                try:
                    size = os.path.getsize(filepath)
                except:
                    size = 0
                rel_path = os.path.relpath(filepath, temp_dir)

                if show_all:
                    files_found.append((rel_path, size))
                else:
                    file_ext = os.path.splitext(filename)[1].lower()
                    if file_ext in filter_exts:
                        files_found.append((rel_path, size))
        return {'files': sorted(files_found), 'header': f'Nested: {archive_file}'}
    except:
        return None
    finally:
        import shutil
        try:
            shutil.rmtree(temp_dir)
        except:
            pass

def list_files_in_archive(zip_path, filter_exts, show_all):
    if zip_path.lower().endswith('.zip'):
        return list_files_in_zip(zip_path, filter_exts, show_all)
    elif zip_path.lower().endswith('.7z'):
        return list_files_in_7z(zip_path, filter_exts, show_all)
    else:
        return {'error': f'Unsupported archive format - {zip_path}', 'files': [], 'header': ''}

def list_files_in_7z(archive_path, filter_exts, show_all):
    import py7zr
    import tempfile
    temp_dir = tempfile.mkdtemp()
    try:
        with py7zr.SevenZipFile(archive_path, 'r') as sz:
            sz.extractall(temp_dir)
        files_found = []
        for root, dirs, files in os.walk(temp_dir):
            for filename in files:
                filepath = os.path.join(root, filename)
                try:
                    size = os.path.getsize(filepath)
                except:
                    size = 0
                rel_path = os.path.relpath(filepath, temp_dir)
                if show_all:
                    files_found.append((rel_path, size))
                else:
                    ext = os.path.splitext(filename)[1].lower()
                    if ext in filter_exts:
                        files_found.append((rel_path, size))
        return {'files': sorted(files_found), 'header': f'Archive: {archive_path}\nTotal files: {len(files_found)}\n' + '=' * 80}
    finally:
        import shutil
        try:
            shutil.rmtree(temp_dir)
        except:
            pass

def list_files_in_zip_subdir(zip_path, sub_path, filter_exts, show_all):
    import tempfile
    temp_dir = tempfile.mkdtemp()
    try:
        with zipfile.ZipFile(zip_path, 'r') as z:
            for member in z.namelist():
                if member.startswith(sub_path + '/') or member == sub_path:
                    z.extract(member, temp_dir)
        extract_subdir = os.path.join(temp_dir, sub_path)
        if os.path.exists(extract_subdir):
            files_found = []
            for root, dirs, files in os.walk(extract_subdir):
                for filename in files:
                    filepath = os.path.join(root, filename)
                    try:
                        size = os.path.getsize(filepath)
                    except:
                        size = 0
                    rel_path = os.path.relpath(filepath, extract_subdir)
                    if show_all:
                        files_found.append((rel_path, size))
                    else:
                        ext = os.path.splitext(filename)[1].lower()
                        if ext in filter_exts:
                            files_found.append((rel_path, size))

            print(f'Archive: {zip_path} / {sub_path}')
            print(f'Total files: {len(files_found)}')
            print('=' * 80)
            for path, size in sorted(files_found):
                print(f'{size:>12,} bytes  {path}')
        else:
            print(f'Error: Subdirectory not found in archive - {sub_path}')
    finally:
        import shutil
        try:
            shutil.rmtree(temp_dir)
        except:
            pass

def list_files_in_dir(dir_path, filter_exts, show_all):
    files_found = []
    for root, dirs, files in os.walk(dir_path):
        for filename in files:
            filepath = os.path.join(root, filename)
            try:
                size = os.path.getsize(filepath)
            except:
                size = 0
            rel_path = os.path.relpath(filepath, dir_path)
            if show_all:
                files_found.append((rel_path, size))
            else:
                ext = os.path.splitext(filename)[1].lower()
                if ext in filter_exts:
                    files_found.append((rel_path, size))
    return {'files': sorted(files_found), 'header': f'Directory: {dir_path}\nTotal files: {len(files_found)}\n' + '=' * 80}

def list_files_in_zip(zip_path, filter_exts, show_all):
    if not os.path.exists(zip_path):
        return {'error': f'File not found - {zip_path}', 'files': [], 'header': ''}

    with zipfile.ZipFile(zip_path, 'r') as z:
        files_found = []
        for name in z.namelist():
            info = z.getinfo(name)
            if info.file_size == 0:
                continue
            if show_all:
                files_found.append((name, info.file_size))
            else:
                ext = os.path.splitext(name)[1].lower()
                if ext in filter_exts:
                    files_found.append((name, info.file_size))
        return {'files': sorted(files_found), 'header': f'Archive: {zip_path}\nTotal files: {len(files_found)}\n' + '=' * 80}

def main():
    import sys
    
    parser = argparse.ArgumentParser(
        description='List files in a zip archive with filtering by type',
        formatter_class=argparse.RawDescriptionHelpFormatter,
        epilog='''
Examples:
  python list_zip_contents.py --target_path archive.zip --all
  python list_zip_contents.py --target_path archive.zip --type java py js
  python list_zip_contents.py --target_path archive.zip --type sql
  python list_zip_contents.py archive.zip --type c cpp h
        ''' 
    )

    parser.add_argument('positional_path', nargs='?', help='Path to the zip archive')
    parser.add_argument('-tp', '--target_path', '--target', dest='target_path', help='Path to the zip archive')
    parser.add_argument('-t', '--type', '--types', dest='types', nargs='+',
                        help=f'File types to filter: {" ".join(EXTENSIONS.keys())}')
    parser.add_argument('-a', '--all', dest='show_all', action='store_true',
                        help='Show all files instead of code files only')
    parser.add_argument('-o', '--output', '--out', dest='output',
                        help='Output file path (e.g., result.csv, result.xlsx)')

    args = parser.parse_args()

    if args.target_path:
        zip_path = args.target_path
    elif args.positional_path:
        zip_path = args.positional_path
    else:
        parser.print_help()
        return

    run_dir = os.getcwd()
    script_dir = os.path.dirname(os.path.abspath(__file__))
    output_dir = run_dir if os.path.exists(run_dir) else script_dir
    
    types = args.types if args.types else list(EXTENSIONS.keys())

    filter_exts = set(EXTENSIONS.values())
    if args.types:
        filter_exts = {EXTENSIONS.get(t, '.' + t) for t in args.types}

    result = list_files_in_path(zip_path, filter_exts, args.show_all)

    if result.get('error'):
        print(result['error'])
        return

    if args.output:
        output_path = args.output
        ext = os.path.splitext(output_path)[1].lower()
        
        target_name = os.path.basename(zip_path)
        target_stem = os.path.splitext(target_name)[0]
        
        if not os.path.isabs(output_path):
            output_path = os.path.join(output_dir, output_path)

    filter_exts = set(EXTENSIONS.values())
    if args.types:
        filter_exts = {EXTENSIONS.get(t, '.' + t) for t in args.types}

    result = list_files_in_path(zip_path, filter_exts, args.show_all)

    if result.get('error'):
        print(result['error'])
        return

    if args.output:
        output_path = args.output
        ext = os.path.splitext(output_path)[1].lower()
        
        target_name = os.path.basename(zip_path)
        target_stem = os.path.splitext(target_name)[0]
        
        if not os.path.isabs(output_path):
            output_path = os.path.join(output_dir, output_path)
        
        if os.path.exists(output_path):
            print(f'File "{output_path}" already exists. Overwrite? (Yes/No): ')
            choice = input().strip().lower()
            if choice != 'yes' and choice != 'y':
                default_name = f'{target_stem}_{ext[1:]}' if ext else target_stem
                print(f'Please modify the file name to: {default_name}')
                new_name = input().strip()
                if new_name:
                    if not os.path.isabs(new_name):
                        new_name = os.path.join(output_dir, new_name)
                    output_path = new_name
                else:
                    output_path = os.path.join(output_dir, default_name)
        
        abs_path = os.path.abspath(output_path)
        if ext == '.csv':
            import csv
            with open(output_path, 'w', newline='', encoding='utf-8') as f:
                writer = csv.writer(f, lineterminator='\n')
                writer.writerow(['Size (bytes)', 'Path'])
                for path, size in result['files']:
                    writer.writerow([size, path])
            print(f'CSV saved to: {abs_path}')
        elif ext in ('.xlsx', '.xls'):
            from openpyxl import Workbook
            wb = Workbook()
            ws = wb.active
            ws.title = 'Files'
            ws.append(['Size (bytes)', 'Path'])
            for path, size in result['files']:
                ws.append([size, path])
            for col in ws.columns:
                max_length = 0
                col_letter = col[0].column_letter
                for cell in col:
                    if cell.value:
                        max_length = max(max_length, len(str(cell.value)))
                ws.column_dimensions[col_letter].width = min(max_length + 2, 80)
            wb.save(output_path)
            print(f'Excel saved: {abs_path}')
        else:
            print(f'Error: Unsupported file extension - {ext}')
            return
    else:
        if result['header']:
            print(result['header'])
        for path, size in result['files']:
            print(f'{size:>12,} bytes  {path}')

if __name__ == '__main__':
    main()

優(yōu)化建議

僅支持csv、xlsx,建議添加md格式,或者結(jié)合格式轉(zhuǎn)換等在線網(wǎng)站。

僅支持后綴,建議添加正則匹配,可以向everything那樣支持任意關(guān)鍵字搜索。

潛在影響,不支持exe格式等文件格式掃描,可能會(huì)被防火墻、EDR等產(chǎn)品或服務(wù)阻止。

到此這篇關(guān)于Python實(shí)現(xiàn)根據(jù)文件后綴統(tǒng)計(jì)文件大小并找出文件位置的文章就介紹到這了,更多相關(guān)Python查找文件大小和位置內(nèi)容請(qǐng)搜索腳本之家以前的文章或繼續(xù)瀏覽下面的相關(guān)文章希望大家以后多多支持腳本之家!

相關(guān)文章

  • python+rsync精確同步指定格式文件

    python+rsync精確同步指定格式文件

    這篇文章主要為大家詳細(xì)介紹了python+rsync精確同步指定格式文件,具有一定的參考價(jià)值,感興趣的小伙伴們可以參考一下
    2019-08-08
  • python如何實(shí)現(xiàn)數(shù)組反轉(zhuǎn)

    python如何實(shí)現(xiàn)數(shù)組反轉(zhuǎn)

    這篇文章主要介紹了python如何實(shí)現(xiàn)數(shù)組反轉(zhuǎn)問(wèn)題,具有很好的參考價(jià)值,希望對(duì)大家有所幫助。如有錯(cuò)誤或未考慮完全的地方,望不吝賜教
    2023-02-02
  • 淺談關(guān)于Python3中venv虛擬環(huán)境

    淺談關(guān)于Python3中venv虛擬環(huán)境

    這篇文章主要介紹了淺談關(guān)于Python3中venv虛擬環(huán)境,小編覺(jué)得挺不錯(cuò)的,現(xiàn)在分享給大家,也給大家做個(gè)參考。一起跟隨小編過(guò)來(lái)看看吧
    2018-08-08
  • 深入理解python中的atexit模塊

    深入理解python中的atexit模塊

    atexit模塊很簡(jiǎn)單,只定義了一個(gè)register函數(shù)用于注冊(cè)程序退出時(shí)的回調(diào)函數(shù),我們可以在這個(gè)回調(diào)函數(shù)中做一些資源清理的操作。下面這篇文章主要介紹了python中atexit模塊的相關(guān)資料,需要的朋友可以參考下。
    2017-03-03
  • 詳解Python的Django框架中manage命令的使用與擴(kuò)展

    詳解Python的Django框架中manage命令的使用與擴(kuò)展

    這篇文章主要介紹了Python的Django框架中manage命令的使用與擴(kuò)展,manage.py使得用戶借助manage命令在命令行中能實(shí)現(xiàn)諸多簡(jiǎn)便的操作,需要的朋友可以參考下
    2016-04-04
  • Python手繪可視化工具cutecharts使用實(shí)例

    Python手繪可視化工具cutecharts使用實(shí)例

    這篇文章主要介紹了Python手繪可視化工具cutecharts使用實(shí)例,文中通過(guò)示例代碼介紹的非常詳細(xì),對(duì)大家的學(xué)習(xí)或者工作具有一定的參考學(xué)習(xí)價(jià)值,需要的朋友可以參考下
    2019-12-12
  • Python中應(yīng)用Winsorize縮尾處理的操作經(jīng)驗(yàn)

    Python中應(yīng)用Winsorize縮尾處理的操作經(jīng)驗(yàn)

    縮尾處理相當(dāng)于對(duì)數(shù)據(jù)進(jìn)行掐頭(尾)去尾,然后再按照一定的方法填補(bǔ)被掐掉的數(shù)據(jù),下面這篇文章主要給給大家介紹了關(guān)于Python中應(yīng)用Winsorize縮尾處理的相關(guān)資料,需要的朋友可以參考下
    2022-07-07
  • Python+PyQt5實(shí)現(xiàn)簡(jiǎn)歷自動(dòng)生成工具

    Python+PyQt5實(shí)現(xiàn)簡(jiǎn)歷自動(dòng)生成工具

    在當(dāng)今競(jìng)爭(zhēng)激烈的求職市場(chǎng)中,一份專(zhuān)業(yè)、規(guī)范的簡(jiǎn)歷是獲得面試機(jī)會(huì)的關(guān)鍵,本文介紹了一個(gè)基于PyQt5的自動(dòng)化簡(jiǎn)歷生成工具,有需要的小伙伴可以參考一下
    2025-05-05
  • Python3如何解決字符編碼問(wèn)題詳解

    Python3如何解決字符編碼問(wèn)題詳解

    字符串是一種數(shù)據(jù)類(lèi)型,但是,字符串比較特殊的是還有一個(gè)編碼問(wèn)題。下面這篇文章主要給大家介紹了關(guān)于Python3如何解決字符編碼問(wèn)題的相關(guān)資料,文中介紹的還是相對(duì)比較詳細(xì)的,需要的朋友可以參考借鑒,下面來(lái)一起看看吧。
    2017-04-04
  • Python 文件操作大全

    Python 文件操作大全

    這篇文章主要介紹了Python 文件操作大全,本文結(jié)合實(shí)例代碼給大家介紹的非常詳細(xì),需要的朋友可以參考下
    2019-09-09

最新評(píng)論

营口市| 东城区| 稷山县| 东光县| 富顺县| 肥东县| 庆阳市| 台东县| 崇仁县| 哈密市| 长治市| 锡林浩特市| 乌拉特后旗| 尉犁县| 梅河口市| 江门市| 田林县| 藁城市| 岳西县| 武夷山市| 蓬安县| 北碚区| 疏勒县| 晋州市| 香格里拉县| 铁力市| 容城县| 会宁县| 陵川县| 叶城县| 东方市| 林州市| 商都县| 庆元县| 塔城市| 佛山市| 五台县| 泽普县| 佛冈县| 河北区| 衡阳县|