
Preventing `abi.encodePacked` Collision Bugs in Solidity
What abi.encodePacked actually does
abi.encodePacked produces a tightly packed byte sequence with minimal padding. For fixed-size values, this is often fine. The problem appears when you combine multiple dynamic types such as string, bytes, bytes[], or arrays of values whose boundaries are not self-describing.
For example, these two inputs produce the same packed bytes:
("ab", "c")("a", "bc")
If you hash both with keccak256(abi.encodePacked(...)), the resulting digest is identical.
That is not a bug in Solidity; it is a consequence of how packed encoding works. The bug appears when developers assume the encoding is uniquely decodable.
Why collisions matter in real contracts
Packed-encoding collisions become dangerous when the resulting hash is used for:
- signature verification
- permit-style approvals
- off-chain order matching
- allowlist or whitelist checks
- commit-reveal schemes
- message authentication
- deterministic IDs for records or orders
A collision means two different logical inputs can map to the same byte sequence. If your contract treats the hash as a unique identifier, an attacker may be able to substitute values that satisfy the same check.
Typical failure pattern
A contract computes a hash like this:
bytes32 digest = keccak256(abi.encodePacked(user, amount, memo));If memo is dynamic, and the contract later compares digest against a signed message or stored commitment, an attacker may be able to craft different (user, amount, memo) combinations that yield the same digest.
The issue is especially subtle because the code looks clean and gas-efficient.
When packed encoding is safe
Packed encoding is not inherently unsafe. It is acceptable when the encoded fields are:
- all fixed-size
- clearly separated by type boundaries
- not user-controlled in a way that creates ambiguity
- not used as a security boundary
Safe examples
These are generally safe:
keccak256(abi.encodePacked(address(this), msg.sender, uint256(123)));
keccak256(abi.encodePacked(bytes32Salt, uint256Id));Why? Because address, bytes32, and uint256 have fixed widths, so the concatenation is unambiguous.
Unsafe examples
These are risky:
keccak256(abi.encodePacked(name, symbol));
keccak256(abi.encodePacked(userInput1, userInput2, userInput3));
keccak256(abi.encodePacked(bytesData, anotherBytesData));If any adjacent values are dynamic, the boundaries may be impossible to reconstruct from the packed bytes alone.
Prefer abi.encode for security-critical hashing
The simplest and most reliable fix is to use abi.encode instead of abi.encodePacked when generating hashes for security-sensitive logic.
abi.encode includes type information and padding, making the encoded output unambiguous.
Comparison
| Encoding method | Ambiguity risk | Typical use case | Security recommendation |
|---|---|---|---|
abi.encode | Low | Hashing, signatures, structured data | Preferred |
abi.encodePacked | Medium to high with dynamic values | Compact byte concatenation | Use only with care |
| Manual concatenation | High | Rarely justified | Avoid |
Example: vulnerable vs. safe message hashing
Consider a simple authorization scheme where a signer approves a withdrawal.
Vulnerable version
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
contract PackedHashVault {
address public signer;
constructor(address _signer) {
signer = _signer;
}
function withdraw(
string calldata recipientTag,
string calldata note,
uint256 amount,
bytes calldata signature
) external {
bytes32 digest = keccak256(
abi.encodePacked(recipientTag, note, amount)
);
address recovered = _recover(digest, signature);
require(recovered == signer, "invalid signature");
// transfer logic omitted
}
function _recover(bytes32 digest, bytes calldata signature)
internal
pure
returns (address)
{
// simplified example
return address(0);
}
}This is dangerous because recipientTag and note are both dynamic strings. Different pairs can produce the same packed bytes.
Safer version
// SPDX-License-Identifier: MIT
pragma solidity ^0.8.24;
contract StructuredHashVault {
address public signer;
constructor(address _signer) {
signer = _signer;
}
function withdraw(
string calldata recipientTag,
string calldata note,
uint256 amount,
bytes calldata signature
) external {
bytes32 digest = keccak256(
abi.encode(recipientTag, note, amount)
);
address recovered = _recover(digest, signature);
require(recovered == signer, "invalid signature");
// transfer logic omitted
}
function _recover(bytes32 digest, bytes calldata signature)
internal
pure
returns (address)
{
// simplified example
return address(0);
}
}This version is much safer because the ABI encoding preserves field boundaries.
Add explicit domain separation
Even with abi.encode, you should avoid hashing raw user data alone. A robust message should include context that makes the digest valid only for one contract, chain, and action.
A good hash often includes:
- a contract-specific domain separator
- the chain ID
- the intended action
- a nonce or unique identifier
- a deadline or expiration
- the structured payload
This prevents a valid message from being reused in another context.
Practical pattern
bytes32 digest = keccak256(
abi.encode(
keccak256("Withdraw(address recipient,uint256 amount,uint256 nonce,uint256 deadline)"),
recipient,
amount,
nonce,
deadline
)
);This is more resilient than concatenating raw strings or bytes.
Use length prefixes when compact encoding is unavoidable
Sometimes you really do want a compact byte string, for example when building a custom commitment scheme or a Merkle leaf. In those cases, you can reduce ambiguity by encoding lengths explicitly.
For example:
bytes memory packed = abi.encodePacked(
uint256(bytes(name).length),
bytes(name),
uint256(bytes(symbol).length),
bytes(symbol)
);
bytes32 leaf = keccak256(packed);By prefixing each dynamic field with its length, you make the encoding self-delimiting.
This is still more error-prone than abi.encode, but it can be appropriate when compactness matters and the format is carefully specified.
Common anti-patterns to avoid
1. Hashing multiple dynamic values with abi.encodePacked
This is the most common mistake. If two or more adjacent fields are dynamic, assume the encoding is ambiguous unless proven otherwise.
2. Using packed hashes as unique IDs without structure
If you store records under keccak256(abi.encodePacked(a, b)), collisions may cause one record to overwrite another or make lookups ambiguous.
3. Mixing user input with protocol metadata without separators
For example:
keccak256(abi.encodePacked(user, action, data));If action and data are dynamic, the boundary between them is not guaranteed.
4. Relying on string concatenation semantics
Human-readable strings are convenient for debugging, but they are poor security primitives. "alice" + "bob" is not a structured message format.
A checklist for safe hashing
Before using a hash in authorization or identity logic, verify the following:
- Are all fields fixed-size?
- If not, are dynamic fields encoded with
abi.encode? - Is the message bound to the current contract address?
- Is the chain ID included?
- Is there a nonce to prevent reuse?
- Is there a deadline if the message is time-sensitive?
- Is the exact encoding documented for off-chain signers?
If you cannot answer these confidently, the hash format is probably too fragile.
Off-chain signing and interoperability
Packed encoding bugs often appear in systems where Solidity contracts interact with backend services, wallets, or indexers. The off-chain code must reproduce the exact same byte layout as the contract.
That means the encoding choice is part of the protocol. If the backend uses one concatenation rule and the contract uses another, signatures will fail or, worse, validate unexpectedly.
Best practices for off-chain compatibility
- Define the message schema in one place.
- Use structured encoding libraries where possible.
- Avoid ad hoc string concatenation.
- Test known vectors from both Solidity and the off-chain implementation.
- Document field order, types, and domain separator rules.
Testing for collision resistance
You do not need to prove cryptographic collision resistance of Keccak-256 itself. Instead, test whether your encoding is ambiguous.
Useful test cases
Try inputs that differ only by boundary placement:
("ab", "c")vs.("a", "bc")("hello", "")vs.("hell", "o")("", "abc")vs.("a", "bc")if additional fields are present
If two distinct logical inputs produce the same packed bytes, your encoding is unsafe for that use case.
Property-based testing idea
In fuzz tests, generate pairs of dynamic strings and assert that your chosen encoding method does not collapse distinct structured inputs into the same hash. If you must use packed encoding, verify that all fields are fixed-size or length-prefixed.
Choosing the right approach
| Scenario | Recommended approach |
|---|---|
| Signature over structured data | abi.encode with domain separation |
| Unique record key from fixed-size fields | abi.encodePacked may be acceptable |
| Commitment over user-supplied strings | abi.encode or length-prefixed format |
| Off-chain order signing | Structured encoding with explicit schema |
| Merkle leaf construction | Prefer a documented, length-safe format |
The key question is not whether packed encoding is shorter. It is whether the encoding is uniquely decodable for the values you are combining.
Practical guidance for production contracts
- Use
abi.encodeby default for security-sensitive hashes. - Treat
abi.encodePackedas a low-level optimization, not a default choice. - Never pack adjacent dynamic values without length prefixes.
- Include contract- and chain-specific domain separation in signed messages.
- Keep the encoding format stable and documented.
- Add tests that intentionally probe ambiguous boundaries.
- Review every hash used for authorization, not just those used for storage keys.
Conclusion
abi.encodePacked is useful, but only when you understand its ambiguity boundaries. The moment a packed hash becomes part of a security decision, you must ensure the encoded message is unambiguous, domain-separated, and reproducible across implementations.
In most cases, abi.encode is the safer default. Use packed encoding only when the fields are fixed-size or when you explicitly add length information. That small discipline prevents a class of bugs that are easy to miss in review and expensive to fix after deployment.
