Hello,

I would like feedback on iv, a small line-oriented utility that
exposes common text transformations directly as command-line
operations rather than through an embedded editing language.

The intended scope is deliberately narrow: substitute, delete,
replace, insert and append; literal or POSIX ERE matching;
stdin/stdout composition; and optional in-place replacement.

Search, count and tail operations are deliberately left to grep,
wc and tail.

This is a design RFC, not a request to merge the current GitHub
tree. The implementation exists to make the proposed interface
and its semantics concrete.

Interface
---------

Some representative operations are:

  iv -s FILE PAT REPL          # first literal match per line
  iv -s FILE PAT REPL -g       # every match per line
  iv -s FILE PAT REPL -E       # POSIX ERE; \1-\9 and &
  iv -d FILE RANGE             # delete lines
  iv -d FILE -m PAT            # delete matching lines
  iv -r FILE RANGE TEXT        # replace lines
  iv -i FILE RANGE TEXT        # insert text

Transformations can also be used as filters:

  iv -s FILE a b --stdout | iv -s - b c --stdout
  iv -s FILE a b --stdout | head -n 1

Ranges are 1-based (1-5, 2-, -3--1).

Text arguments are literal. "-" means stdin. An existing path is
not treated as a file to read: iv -r foo 3 bar must keep meaning
"bar" if a file named bar appears in the working directory. File
contents are supplied by redirection:

  iv -p dest - < snippet.c

The distinction I am interested in is not whether these
transformations can already be expressed with sed or awk. They
can.

The proposed abstraction is that a common transformation is
represented directly by argv rather than by a small editing
program. iv intentionally does not grow sed's programming model:
no branching, hold space, addresses or script files.

Likewise, operations already represented well by grep, wc and
tail are outside its scope and compose through pipes.

Text model
----------

iv is deliberately byte-oriented.

  LC_ALL=C
  line = bytes through '\n', or through EOF
  POSIX ERE under -E
  invalid UTF-8 is data
  NUL input is rejected

-F is a single-byte field delimiter, not CSV parsing.

These are interface choices rather than accidental implementation
limitations.

In-place replacement and backups
--------------------------------

In-place editing uses a temporary file in the destination
directory:

  transform
    -> exclusive openat() temporary
    -> checked write
    -> fsync()
    -> verify destination inode
    -> renameat()

A failed transformation or write before the rename leaves the
original pathname referring to the original contents.

Backups are optional and use the existing GNU Coreutils
convention:

  -b
  --backup[=METHOD]
  -S SUFFIX
  VERSION_CONTROL
  SIMPLE_BACKUP_SUFFIX

with none, numbered, existing and simple semantics.

No backup is created unless requested.

Some filesystem semantics are intentionally explicit:

  * regular files only;
  * symlinks edit the referent without replacing the symlink;
  * dangling symlinks are rejected;
  * replacing one hardlink pathname does not modify the other
    links;
  * mode and uid/gid are retained when permitted;
  * xattrs and ACLs are currently not preserved;
  * SIGINT/SIGTERM/SIGHUP clean an outstanding temporary;
    SIGKILL may leave it behind;
  * this is not intended as a security boundary for directories
    writable by an untrusted party.

EPIPE on stdout is treated as successful early termination, so
e.g.

  iv -s large-file a b --stdout | head -n 1

exits without a Broken pipe diagnostic.

Implementation evidence
-----------------------

Performance is not the reason for proposing the interface. I did,
however, want to establish that the narrower model does not
impose a performance or memory penalty compared with the general
tools used for the same operations.

Current implementation:

  C, libc only at runtime
  -O2
  make test
  make test-musl

The test suites currently pass under both glibc and musl, in
addition to smoke/safety tests. Cases include backup methods,
ENOSPC, EINTR, EPIPE, signals, concurrent pathname replacement,
dangling symlinks, hardlinks, missing final newline, huge lines,
empty EREs, literal text vs stdin, and metadata behavior.

The benchmark harness first requires byte-identical output for
every declared-equivalent pair, then runs 11 interleaved trials
after one warmup with CPU affinity. RSS is the median GNU time
peak RSS.

On my x86_64 Linux test host, against GNU sed 4.10 and
mawk 1.3.4:

                             1M lines       3M lines
literal substitution
  iv                         159 ms          470 ms
  GNU sed 4.10               495 ms         1477 ms
  mawk                       215 ms          635 ms

ERE substitution
  iv                        1506 ms         4712 ms
  GNU sed 4.10              2552 ms         7968 ms

delete matching lines
  iv                         103 ms          293 ms
  GNU sed 4.10               197 ms          586 ms
  mawk                       116 ms          329 ms

Median peak RSS for the stdout jobs is about 1.6 MB for iv and
2.6-2.7 MB for sed/mawk on this host.

These numbers are not intended to claim that iv replaces either
tool. mawk in particular is included as a useful reference for
efficient one-pass processing. The claim is only that the
explicit interface can be implemented competitively while
retaining bounded memory and the filesystem semantics above.

The harness, corpus generation and equivalence checks are
reproducible from the tree.

Scope
-----

The current implementation has more convenience around the core
operations, but the proposed utility is intentionally not a new
text-processing language.

In particular, I am not proposing replacements for grep, wc or
tail, nor sed-style branching/hold-space programming, nor
locale-aware column processing.

The design question I would like feedback on is:

Is an explicit-operation interface for common line-oriented
transformations sufficiently distinct and useful to justify a
separate GNU utility?

If so, I would also like feedback on whether Coreutils is an
appropriate home for it and which parts of the current interface
should be reduced before proposing a patch.

If not, I would be particularly interested in whether the
objection is to the abstraction itself, its current scope, or
specifically to Coreutils as its home.

Current implementation:

  https://github.com/IRodriguez13/iv

GPL-3.0-or-later.

Thanks,

Iván Ezequiel Rodriguez
[email protected]

Reply via email to