最新国产好看的视频,伊人天堂AV在线,国产Aaaaaa视频,蜜臀视频在线观看一区,人妻av色图,密臀久久久精品影片,青青视频免费观看毛片,久草在线观看视,国产三级精品色情在线

python實現將html表格轉換成CSV文件的方法

 更新時間:2015年06月28日 14:53:46   作者:秋風秋雨  
這篇文章主要介紹了python實現將html表格轉換成CSV文件的方法,涉及Python操作csv文件的相關技巧,需要的朋友可以參考下

本文實例講述了python實現將html表格轉換成CSV文件的方法。分享給大家供大家參考。具體如下:

使用方法:python html2csv.py *.html
這段代碼使用了 HTMLParser 模塊

#!/usr/bin/python
# -*- coding: iso-8859-1 -*-
# Hello, this program is written in Python - http://python.org
programname = 'html2csv - version 2002-09-20 - http://sebsauvage.net'
import sys, getopt, os.path, glob, HTMLParser, re
try:  import psyco ; psyco.jit() # If present, use psyco to accelerate the program
except: pass
def usage(progname):
  ''' Display program usage. '''
  progname = os.path.split(progname)[1]
  if os.path.splitext(progname)[1] in ['.py','.pyc']: progname = 'python '+progname
  return '''%s
A coarse HTML tables to CSV (Comma-Separated Values) converter.
Syntax  : %s source.html
Arguments : source.html is the HTML file you want to convert to CSV.
      By default, the file will be converted to csv with the same
      name and the csv extension (source.html -> source.csv)
      You can use * and ?.
Examples  : %s mypage.html
      : %s *.html
This program is public domain.
Author : Sebastien SAUVAGE <sebsauvage at sebsauvage dot net>
     http://sebsauvage.net
''' % (programname, progname, progname, progname)
class html2csv(HTMLParser.HTMLParser):
  ''' A basic parser which converts HTML tables into CSV.
    Feed HTML with feed(). Get CSV with getCSV(). (See example below.)
    All tables in HTML will be converted to CSV (in the order they occur
    in the HTML file).
    You can process very large HTML files by feeding this class with chunks
    of html while getting chunks of CSV by calling getCSV().
    Should handle badly formated html (missing <tr>, </tr>, </td>,
    extraneous </td>, </tr>...).
    This parser uses HTMLParser from the HTMLParser module,
    not HTMLParser from the htmllib module.
    Example: parser = html2csv()
         parser.feed( open('mypage.html','rb').read() )
         open('mytables.csv','w+b').write( parser.getCSV() )
    This class is public domain.
    Author: Sébastien SAUVAGE <sebsauvage at sebsauvage dot net>
        http://sebsauvage.net
    Versions:
      2002-09-19 : - First version
      2002-09-20 : - now uses HTMLParser.HTMLParser instead of htmllib.HTMLParser.
            - now parses command-line.
    To do:
      - handle <PRE> tags
      - convert html entities (&name; and &#ref;) to Ascii.
      '''
  def __init__(self):
    HTMLParser.HTMLParser.__init__(self)
    self.CSV = ''   # The CSV data
    self.CSVrow = ''  # The current CSV row beeing constructed from HTML
    self.inTD = 0   # Used to track if we are inside or outside a <TD>...</TD> tag.
    self.inTR = 0   # Used to track if we are inside or outside a <TR>...</TR> tag.
    self.re_multiplespaces = re.compile('\s+') # regular expression used to remove spaces in excess
    self.rowCount = 0 # CSV output line counter.
  def handle_starttag(self, tag, attrs):
    if  tag == 'tr': self.start_tr()
    elif tag == 'td': self.start_td()
  def handle_endtag(self, tag):
    if  tag == 'tr': self.end_tr()
    elif tag == 'td': self.end_td()     
  def start_tr(self):
    if self.inTR: self.end_tr() # <TR> implies </TR>
    self.inTR = 1
  def end_tr(self):
    if self.inTD: self.end_td() # </TR> implies </TD>
    self.inTR = 0      
    if len(self.CSVrow) > 0:
      self.CSV += self.CSVrow[:-1]
      self.CSVrow = ''
    self.CSV += '\n'
    self.rowCount += 1
  def start_td(self):
    if not self.inTR: self.start_tr() # <TD> implies <TR>
    self.CSVrow += '"'
    self.inTD = 1
  def end_td(self):
    if self.inTD:
      self.CSVrow += '",' 
      self.inTD = 0
  def handle_data(self, data):
    if self.inTD:
      self.CSVrow += self.re_multiplespaces.sub(' ',data.replace('\t',' ').replace('\n','').replace('\r','').replace('"','""'))
  def getCSV(self,purge=False):
    ''' Get output CSV.
      If purge is true, getCSV() will return all remaining data,
      even if <td> or <tr> are not properly closed.
      (You would typically call getCSV with purge=True when you do not have
      any more HTML to feed and you suspect dirty HTML (unclosed tags). '''
    if purge and self.inTR: self.end_tr() # This will also end_td and append last CSV row to output CSV.
    dataout = self.CSV[:]
    self.CSV = ''
    return dataout
if __name__ == "__main__":
  try: # Put getopt in place for future usage.
    opts, args = getopt.getopt(sys.argv[1:],None)
  except getopt.GetoptError:
    print usage(sys.argv[0]) # print help information and exit:
    sys.exit(2)
  if len(args) == 0:
    print usage(sys.argv[0]) # print help information and exit:
    sys.exit(2)    
  print programname
  html_files = glob.glob(args[0])
  for htmlfilename in html_files:
    outputfilename = os.path.splitext(htmlfilename)[0]+'.csv'
    parser = html2csv()
    print 'Reading %s, writing %s...' % (htmlfilename, outputfilename)
    try:
      htmlfile = open(htmlfilename, 'rb')
      csvfile = open( outputfilename, 'w+b')
      data = htmlfile.read(8192)
      while data:
        parser.feed( data )
        csvfile.write( parser.getCSV() )
        sys.stdout.write('%d CSV rows written.\r' % parser.rowCount)
        data = htmlfile.read(8192)
      csvfile.write( parser.getCSV(True) )
      csvfile.close()
      htmlfile.close()
    except:
      print 'Error converting %s    ' % htmlfilename
      try:  htmlfile.close()
      except: pass
      try:  csvfile.close()
      except: pass
  print 'All done. '

希望本文所述對大家的Python程序設計有所幫助。

相關文章

  • Python代碼實現找到列表中的奇偶異常項

    Python代碼實現找到列表中的奇偶異常項

    這篇文章主要介紹了Python代碼實現找到列表中的奇偶異常項,文章內容主要利用Python代碼實現了從輸入列表中尋找奇偶異常項,需要的朋友可以參考一下
    2021-11-11
  • python使用xlrd模塊讀取xlsx文件中的ip方法

    python使用xlrd模塊讀取xlsx文件中的ip方法

    今天小編就為大家分享一篇python使用xlrd模塊讀取xlsx文件中的ip方法,具有很好的參考價值,希望對大家有所幫助。一起跟隨小編過來看看吧
    2019-01-01
  • numpy給array增加維度np.newaxis的實例

    numpy給array增加維度np.newaxis的實例

    今天小編就為大家分享一篇numpy給array增加維度np.newaxis的實例,具有很好的參考價值,希望對大家有所幫助。一起跟隨小編過來看看吧
    2018-11-11
  • python3中數組逆序輸出方法

    python3中數組逆序輸出方法

    在本篇文章里小編給大家整理的是一篇關于python3中數組逆序輸出方法內容,有需要的朋友們可以學習下。
    2020-12-12
  • Python安裝lz4-0.10.1遇到的坑

    Python安裝lz4-0.10.1遇到的坑

    本篇文章給大家分享了Python安裝lz4-0.10.1的詳細過程以及遇到的坑,需要的讀者們參考下。
    2018-05-05
  • Python獲取、格式化當前時間日期的方法

    Python獲取、格式化當前時間日期的方法

    在本篇文章里小編給大家整理的是關于Python獲取、格式化當前時間日期的方法,對此有需要的朋友們可以學習參考下。
    2020-02-02
  • django如何自己創(chuàng)建一個中間件

    django如何自己創(chuàng)建一個中間件

    這篇文章主要介紹了django如何自己創(chuàng)建一個中間件,文中通過示例代碼介紹的非常詳細,對大家的學習或者工作具有一定的參考學習價值,需要的朋友可以參考下
    2019-07-07
  • python中l(wèi)ower函數實現方法及用法講解

    python中l(wèi)ower函數實現方法及用法講解

    在本篇文章里小編給大家整理的是一篇關于python中l(wèi)ower函數實現方法及用法講解內容,有需要的朋友們可以學習參考下。
    2020-12-12
  • 如何利用Python連接MySQL數據庫實現數據儲存

    如何利用Python連接MySQL數據庫實現數據儲存

    當我們學習了mysql數據庫后,我們會想著該如何將python和mysql結合起來運用,下面這篇文章主要給大家介紹了關于如何利用Python連接MySQL數據庫實現數據儲存的相關資料,需要的朋友可以參考下
    2021-11-11
  • 利用Python輕松解析XML文件

    利用Python輕松解析XML文件

    XML文件在數據處理和配置存儲中非常常見,但手動解析它們可能會讓人頭疼,Python提供了多種簡單高效的方法來處理XML文件,下面小編就來和大家詳細介紹一下吧
    2025-04-04

最新評論

静乐县| 西宁市| 故城县| 灵川县| 岚皋县| 德庆县| 乐平市| 扎鲁特旗| 镇安县| 莱州市| 高台县| 八宿县| 深州市| 北票市| 婺源县| 会理县| 凉山| 通山县| 镇江市| 株洲县| 连州市| 三台县| 仁寿县| 浏阳市| 庄浪县| 芮城县| 民丰县| 三亚市| 抚顺市| 台东县| 昆明市| 康定县| 安达市| 南开区| 丹江口市| 漳平市| 屏东县| 景德镇市| 上饶县| 白山市| 安义县|