LALR Грамматика для преобразования текста в CSV ⇐ Python

Программы на Python
Anonymous
LALR Грамматика для преобразования текста в CSV

Сообщение Anonymous »

У меня есть выходные данные трассировки процессора, имеющие следующий формат:

Код: Выделить всё

Time    Cycle   PC  Instr   Decoded instruction Register and memory contents
905ns              86 00000e36 00a005b3 c.add            x11,  x0, x10       x11=00000e5c x10:00000e5c
915ns              87 00000e38 00000693 c.addi           x13,  x0, 0         x13=00000000
925ns              88 00000e3a 00000613 c.addi           x12,  x0, 0         x12=00000000
935ns              89 00000e3c 00000513 c.addi           x10,  x0, 0         x10=00000000
945ns              90 00000e3e 2b40006f c.jal             x0, 692
975ns              93 000010f2 0d01a703 lw               x14, 208(x3)        x14=00002b20  x3:00003288  PA:00003358
985ns              94 000010f6 00a00333 c.add             x6,  x0, x10        x6=00000000 x10:00000000
995ns              95 000010f8 14872783 lw               x15, 328(x14)       x15=00000000 x14:00002b20  PA:00002c68
1015ns              97 000010fc 00079563 c.bne            x15,  x0, 10        x15:00000000
Предположительно, это разделитель \t, однако это не так, поскольку кое-где встречаются строчные пробелы. Я хочу преобразовать это в формат .csv со строкой заголовка и последующими записями. Например:

Код: Выделить всё

Time,Cycle,PC,Instr,Decoded instruction,Register and memory contents
905ns,86,00000e36,00a005b3,"c.add x11, x0, x10", x11=00000e5c x10:00000e5c
915ns,87,00000e38,00000693,"c.addi x13, x0, 0", x13=00000000
...
Для этого я использую Lark в Python3 (>=3.10). И я придумал следующую грамматику для исходного формата:
Lark Grammar

Код: Выделить всё

start: header NEWLINE entries+

# Header is expected to be
# Time\tCycle\tPC\tInstr\tDecoded instruction\tRegister and memory contents
header: HEADER_FIELD+

# Entries are expected to be e.g.,
#     85ns               4 00000180 00003197 auipc             x3, 0x3000          x3=00003180
entries: TIME                \
CYCLE               \
PC                  \
INSTR               \
DECODED_INSTRUCTION \
reg_and_mem? NEWLINE

reg_and_mem: REG_AND_MEM+

///////////////
// TERMINALS //
///////////////

HEADER_FIELD: /
[a-z ]+  # Characters that are optionally separated by a single space
/xi

TIME: /
[\d\.]+    # One or more digits
[smunp]s   # Time unit
/x

CYCLE: INT

PC: HEXDIGIT+

INSTR: HEXDIGIT+

DECODED_INSTRUCTION: /
[a-z\.]+             # Instruction mnemonic
([-a-z0-9, ()]+)?    # Optional operand part (rd,rs1,rs2, etc.)
(?=                  # Stop when
x[0-9]{1,2}[=:]  # Either you hit an xN= or xN:
|PA:             # or you meet PA:
|\s+$            # or there is no REG_AND_MEM and you meet a \n
)
/xi

REG_AND_MEM: /
(?:[x[0-9]+|PA)
[=|:]
[0-9a-f]+
/xi

///////////////
// IMPORTS   //
///////////////

%import common.HEXDIGIT
%import common.NUMBER
%import common.INT
%import common.UCASE_LETTER
%import common.CNAME
%import common.NUMBER
%import common.WS_INLINE
%import common.WS
%import common.NEWLINE

///////////////
// IGNORE    //
///////////////

%ignore WS_INLINE
Вот мой простой код драйвера:

Код: Выделить всё

import lark

class TraceTransformer(lark.Transformer):

def start(self, args):
return lark.Discard

def header(self, fields):

return [str(field) for field in fields]

def entries(self, args):
print(args)
...

# the grammar provided above
# stored in the same directory
# as this file
parser = lark.Lark(grammar=open("grammar.lark").read(),
start="start",
parser="lalr",
transformer=TraceTransformer())

# This is parsed by the grammar without problems
# Note that I omit from  the  c.addi the operand
# part and its still parsed.  This is ok as  some
# mnemonics do not have operands  (e.g., fence).
dummy_text_ok1 = r"""Time    Cycle   PC  Instr   Decoded instruction Register and memory contents
905ns              86 00000e36 00a005b3 c.add            x11,  x0, x10       x11=00000e5c x10:00000e5c
915ns              87 00000e38 00000693 c.addi           x13,  x0, 0         x13=00000000
925ns              88 00000e3a 00000613 c.addi                  x12=00000000
935ns              89 00000e3c 00000513 c.addi           x10,  x0, 0         x10=00000000"""

# Now here starts trouble. Note that here we don't
# have a REG_AND_MEM part on the jump instruction.
# However this is still parsed with no errors.
dummy_text_ok2 = r"""Time    Cycle   PC  Instr   Decoded instruction Register and memory     945ns              90 00000e3e 2b40006f c.jal             x0, 692       """

# But here, when the parser meets the line of cjal
# where there is no REG_AND_MEM part and a  follow
# up entry exists we have an issue.
dummy_text_problematic = r"""Time    Cycle   PC  Instr   Decoded instruction Register and memory contents
905ns              86 00000e36 00a005b3 c.add            x11,  x0, x10       x11=00000e5c x10:00000e5c
915ns              87 00000e38 00000693 c.addi           x13,  x0, 0         x13=00000000
925ns              88 00000e3a 00000613 c.addi           x12,  x0, 0         x12=00000000
935ns              89 00000e3c 00000513 c.addi           x10,  x0, 0         x10=00000000
945ns              90 00000e3e 2b40006f c.jal             x0, 692
975ns              93 000010f2 0d01a703 lw               x14, 208(x3)        x14=00002b20  x3:00003288  PA:00003358
985ns              94 000010f6 00a00333 c.add             x6,  x0, x10        x6=00000000 x10:00000000
995ns              95 000010f8 14872783 lw               x15, 328(x14)       x15=00000000 x14:00002b20  PA:00002c68
1015ns              97 000010fc 00079563 c.bne            x15,  x0, 10        x15:00000000
"""

parser.parse(dummy_text_ok1)
parser.parse(dummy_text_ok2)
parser.parse(dummy_text_problematic)
Ошибка выполнения

Код: Выделить всё

No terminal matches 'c' in the current parser context, at line 6 col 45

945ns              90 00000e3e 2b40006f c.jal             x0, 692
^
Expected one of:
* DECODED_INSTRUCTION
Итак, это означает, что правило DECODED_INSTRUCTION работает не так, как ожидалось.
Правило

Код: Выделить всё

DECODED_INSTRUCTION: /
[a-z\.]+             # Instruction mnemonic
([-a-z0-9, ()]+)?    # Optional operand part (rd,rs1,rs2, etc.)
(?=                  # Stop when
x[0-9]{1,2}[=:]  # Either you hit an xN= or xN:
|PA:             # or you meet PA:
|\s+$            # or there is no REG_AND_MEM and you meet a \n
)
/xi
Это правило действительно сложное, оно должно соответствовать всей ISA процессора, который, кстати, находится в RISC-V. Итак, шаг за шагом у меня есть
  • Регулярное выражение мнемоники инструкции в виде последовательности символов AZ и необязательных точек (.)
  • Необязательная часть операнда (в ISA существуют инструкции без операндов).
Это было сложно. Вместо того, чтобы учитывать все возможные варианты инструкций в моих правилах выше, я решил использовать тот факт, что в следующем столбце (содержимое регистра и памяти) существуют символы, которых нет ни в одном варианте инструкций ISA. Здесь на помощь приходит упреждающая часть регулярного выражения. Я останавливаюсь, когда
  • Либо я достиг части xN=, либо части xN: поля
  • Либо я дошел до ПА: часть поля
  • ИЛИ дошел до конца строки (), поскольку поле не существует.
Однако последний случай, похоже, не работает должным образом, как показано в приведенном выше примере. . На мой взгляд, нормально либо остановиться, когда вы соответствуете одному из двух критериев, ИЛИ вы столкнулись с новой строкой (подразумевается, что следующая часть опущена для текущей записи). Я что-то ошибся в регулярном выражении?

Подробнее здесь: https://stackoverflow.com/questions/790 ... ext-to-csv

Вернуться в «Python»