Human vs. AI – Diff-based line-level provenance for text under agentic editing
47 points by eighttrigrams 15 hours ago | 13 comments
spuz 10 hours ago
> A git repository is already a history of versions each carrying a provenance marker — every revision of the file, in order, with the author of the change that made it.
replyMaybe I'm missing how people use AI these days but when I have an agent working locally, all git commits have my authorship attached.
andai 3 hours ago
Claude's system prompt tells it to sign git commits with "co-authored by Claude [model]".
reply(Codex doesn't do that.)
GitHub renders this as two authors for those commits.
eighttrigrams 10 hours ago
I take care of that in two ways, 1) is via claude hooks, I make sure these are tagged as claude, 2) my sandboxed agents have git env vars set, with the same effect.
replyspuz 10 hours ago
Does Claude run git commit on every change it makes? Doesn't that pollute your git history?
replyeighttrigrams 10 hours ago
Although that would be an option (I've considered commit on every edit at one point), what I actually ended up doing is that I hook into all Bash tool calls, look if there's a git commit in there, and if so, I prepend it with the right git env variables.
replydrusepth 9 hours ago
A /precommit tool can be very powerful -- not only can it run the tests, do light code review on its own, spin up sub-agents to review performance/architecture/etc impacts, but it can also write hilariously-verbose commit messages once everything passes, detailing not only what changed but why it changed.. tagged with Claude as a co-author.
replyalansaber 11 hours ago
Cool, we do the same thing, but we also denote when a line is "AI generated but was modified by a human" (aka human made anything upwards of a 1 character change).
replyeighttrigrams 10 hours ago
I'll probably be annoyed enough soon that I want that level of granularity, too
reply
> Text a human wrote or edited should be considered close to sacred: an agent should be hesitant and have a very good reason to touch it.
Not sure I agree (particularly with code, not prose). If code is risky to change for reasons that aren't obvious, it should be commented as such. It doesn't matter how the bytes were generated.
> Another use case: the README.md, originally generated, where you rewrite the opening paragraphs. The agent should feel free to redo or append parts further downwards but should really think twice changing anything in the opener.
This, I understand.
For coding, this is more useful for heavily vibecoded projects. Thb, for my own projects most likely I will end up using it only for specs and documentation, maybe in rare cases for some tests. But, I use it for agentic memory, where I can edit the memories as well, and the agent should tend to preserve my edits when updating them.
In general, my main use case is tracking texts in knowledge management systems.