diff --git a/README.md b/README.md index 38400d4..12ae940 100644 --- a/README.md +++ b/README.md @@ -15,11 +15,12 @@ workflows, especially when dealing with multiple audio inputs or outputs. - [6. Audio Channel Conv and Resampler](#6-audio-channel-conv-and-resampler) - [7. Audio Information](#7-audio-information) - [8. Audio Cut](#8-audio-cut) - - [9. Audio Blend](#9-audio-blend) - - [10. Audio Test Signal Generator](#10-audio-test-signal-generator) - - [11. Audio Musical Note](#11-audio-musical-note) - - [12. Audio Join 2 Channels](#12-audio-join-2-channels) - - [13. Audio Split 2 Channels](#13-audio-split-2-channels) + - [9. Audio Concatenate](#8-audio-concatenate) + - [10. Audio Blend](#9-audio-blend) + - [11. Audio Test Signal Generator](#10-audio-test-signal-generator) + - [12. Audio Musical Note](#11-audio-musical-note) + - [13. Audio Join 2 Channels](#12-audio-join-2-channels) + - [14. Audio Split 2 Channels](#13-audio-split-2-channels) - [🚀 Installation](#-installation) - [📦 Dependencies](#-dependencies) - [🖼️ Examples](#️-examples) @@ -149,7 +150,21 @@ workflows, especially when dealing with multiple audio inputs or outputs. - **Output:** - `audio_out` (AUDIO): The selected portion of the audio. -### 9. Audio Blend +### 9. Audio Concatenate + - **Display Name:** `Audio Concatenate` + - **Internal Name:** `SET_AudioConcatenate` + - **Category:** `audio/manipulation` + - **Description:** Joins two audio clips end-to-end in time. This is the counterpart to the "Audio Cut" node. + - **Inputs:** + - `audio1` (AUDIO): The first audio clip to appear in the sequence. + - `audio2` (AUDIO): The second audio clip to be appended to the end of the first. + - **Output:** + - `audio_out` (AUDIO): A single audio clip containing `audio1` followed immediately by `audio2`. + - **Behavior Details:** + - **Alignment:** Before concatenation, the two audio inputs are aligned to have the same sample rate and channel count, using the same logic as the "Batch Audios" node. + - **Batch Handling:** If the inputs have different batch sizes, the last item of the shorter batch is repeated to ensure the output batch size matches the larger of the two inputs. + +### 10. Audio Blend - **Display Name:** `Audio Blend` - **Internal Name:** `SET_AudioBlend` - **Category:** `audio/manipulation` @@ -162,7 +177,7 @@ workflows, especially when dealing with multiple audio inputs or outputs. - **Output:** - `audio_out` (AUDIO): The blended audio. -### 10. Audio Test Signal Generator +### 11. Audio Test Signal Generator - **Display Name:** `Audio Test Signal Generator` - **Internal Name:** `SET_AudioTestSignalGenerator` - **Category:** `audio/generation` @@ -186,7 +201,7 @@ workflows, especially when dealing with multiple audio inputs or outputs. - **Output:** - `audio_out` (AUDIO): The generated test signal. -### 11. Audio Musical Note +### 12. Audio Musical Note - **Display Name:** `Audio Musical Note` - **Internal Name:** `SET_AudioMusicalNote` - **Category:** `audio/generation` @@ -197,7 +212,7 @@ workflows, especially when dealing with multiple audio inputs or outputs. - **Output:** - `frequency` (FLOAT): The calculated frequency of the note in Hz. -### 12. Audio Join 2 Channels +### 13. Audio Join 2 Channels - **Display Name:** `Audio Join 2 Channels` - **Internal Name:** `SET_AudioJoin2Channels` - **Category:** `audio/manipulation` @@ -212,7 +227,7 @@ workflows, especially when dealing with multiple audio inputs or outputs. - **Alignment:** The two mono signals are then aligned to have the same sample rate and length, using the same logic as the "Batch Audios" node (resamples to match `audio_left`'s SR, pads to match the longest duration). - **Batch Handling:** If the inputs have different batch sizes, the last item of the shorter batch is repeated to match the length of the longer batch. -### 13. Audio Split 2 Channels +### 14. Audio Split 2 Channels - **Display Name:** `Audio Split 2 Channels` - **Internal Name:** `SET_AudioSplit2Channels` - **Category:** `audio/manipulation` @@ -254,6 +269,7 @@ Once installed the examples are available in the ComfyUI workflow templates, in and the sample rate. - [generate_and_blend.json](example_workflows/generate_and_blend.json): Shows how to generate four musical notes and blend them together to create a chord. +- [cut_and_concat.json](example_workflows/cut_and_concat.json): Shows how to cut and concatenate audio. ## 📝 Usage Notes diff --git a/example_workflows/cut_and_concat.jpg b/example_workflows/cut_and_concat.jpg new file mode 100644 index 0000000..f2d843a Binary files /dev/null and b/example_workflows/cut_and_concat.jpg differ diff --git a/example_workflows/cut_and_concat.json b/example_workflows/cut_and_concat.json new file mode 100644 index 0000000..11cabdb --- /dev/null +++ b/example_workflows/cut_and_concat.json @@ -0,0 +1 @@ +{"id":"6272cce7-7157-4505-8a62-76cf2ed02bdb","revision":0,"last_node_id":11,"last_link_id":7,"nodes":[{"id":3,"type":"SET_AudioCut","pos":[1559.5836181640625,1503.6956787109375],"size":[270,82],"flags":{},"order":6,"mode":0,"inputs":[{"localized_name":"audio","name":"audio","type":"AUDIO","link":2},{"localized_name":"start_time","name":"start_time","type":"STRING","widget":{"name":"start_time"},"link":null},{"localized_name":"end_time","name":"end_time","type":"STRING","widget":{"name":"end_time"},"link":null}],"outputs":[{"localized_name":"audio_out","name":"audio_out","type":"AUDIO","links":[4,7]}],"properties":{"aux_id":"set-soft/ComfyUI-AudioBatch","ver":"c5e110e58b6bb0f5c777084a1edcadb7b86eb5b7","Node name for S&R":"SET_AudioCut"},"widgets_values":["2","3"],"color":"#323","bgcolor":"#535"},{"id":1,"type":"SET_AudioTestSignalGenerator","pos":[1202.8035888671875,1446.704833984375],"size":[275.0025329589844,322],"flags":{},"order":0,"mode":0,"inputs":[{"localized_name":"waveform_type","name":"waveform_type","type":"COMBO","widget":{"name":"waveform_type"},"link":null},{"localized_name":"frequency","name":"frequency","type":"FLOAT","widget":{"name":"frequency"},"link":null},{"localized_name":"frequency_end","name":"frequency_end","type":"FLOAT","widget":{"name":"frequency_end"},"link":null},{"localized_name":"amplitude","name":"amplitude","type":"FLOAT","widget":{"name":"amplitude"},"link":null},{"localized_name":"dc_offset","name":"dc_offset","type":"FLOAT","widget":{"name":"dc_offset"},"link":null},{"localized_name":"phase","name":"phase","type":"FLOAT","widget":{"name":"phase"},"link":null},{"localized_name":"duration","name":"duration","type":"STRING","widget":{"name":"duration"},"link":null},{"localized_name":"sample_rate","name":"sample_rate","type":"INT","widget":{"name":"sample_rate"},"link":null},{"localized_name":"batch_size","name":"batch_size","type":"INT","widget":{"name":"batch_size"},"link":null},{"localized_name":"channels","name":"channels","type":"INT","widget":{"name":"channels"},"link":null},{"localized_name":"seed","name":"seed","shape":7,"type":"INT","widget":{"name":"seed"},"link":null}],"outputs":[{"localized_name":"audio_out","name":"audio_out","type":"AUDIO","links":[1,2]}],"properties":{"aux_id":"set-soft/ComfyUI-AudioBatch","ver":"c5e110e58b6bb0f5c777084a1edcadb7b86eb5b7","Node name for S&R":"SET_AudioTestSignalGenerator"},"widgets_values":["sweep",440,880,0.5,0,0,"3.0",44100,1,1,168738277895224,"randomize"],"color":"#232","bgcolor":"#353"},{"id":7,"type":"PreviewAudio","pos":[1897.85888671875,1574.5079345703125],"size":[270,88],"flags":{},"order":9,"mode":0,"inputs":[{"localized_name":"audio","name":"audio","type":"AUDIO","link":7},{"localized_name":"audioUI","name":"audioUI","type":"AUDIO_UI","widget":{"name":"audioUI"},"link":null}],"outputs":[],"properties":{"cnr_id":"comfy-core","ver":"0.3.43","Node name for S&R":"PreviewAudio"},"widgets_values":[],"color":"#222","bgcolor":"#000"},{"id":6,"type":"PreviewAudio","pos":[1897.85888671875,1260.73876953125],"size":[270,88],"flags":{},"order":7,"mode":0,"inputs":[{"localized_name":"audio","name":"audio","type":"AUDIO","link":6},{"localized_name":"audioUI","name":"audioUI","type":"AUDIO_UI","widget":{"name":"audioUI"},"link":null}],"outputs":[],"properties":{"cnr_id":"comfy-core","ver":"0.3.43","Node name for S&R":"PreviewAudio"},"widgets_values":[],"color":"#222","bgcolor":"#000"},{"id":5,"type":"PreviewAudio","pos":[2226.381591796875,1435.8328857421875],"size":[270,88],"flags":{},"order":10,"mode":0,"inputs":[{"localized_name":"audio","name":"audio","type":"AUDIO","link":5},{"localized_name":"audioUI","name":"audioUI","type":"AUDIO_UI","widget":{"name":"audioUI"},"link":null}],"outputs":[],"properties":{"cnr_id":"comfy-core","ver":"0.3.43","Node name for S&R":"PreviewAudio"},"widgets_values":[],"color":"#222","bgcolor":"#000"},{"id":2,"type":"SET_AudioCut","pos":[1559.5836181640625,1364.7237548828125],"size":[270,82],"flags":{},"order":5,"mode":0,"inputs":[{"localized_name":"audio","name":"audio","type":"AUDIO","link":1},{"localized_name":"start_time","name":"start_time","type":"STRING","widget":{"name":"start_time"},"link":null},{"localized_name":"end_time","name":"end_time","type":"STRING","widget":{"name":"end_time"},"link":null}],"outputs":[{"localized_name":"audio_out","name":"audio_out","type":"AUDIO","links":[3,6]}],"properties":{"aux_id":"set-soft/ComfyUI-AudioBatch","ver":"c5e110e58b6bb0f5c777084a1edcadb7b86eb5b7","Node name for S&R":"SET_AudioCut"},"widgets_values":["0","1"],"color":"#323","bgcolor":"#535"},{"id":4,"type":"SET_AudioConcatenate","pos":[1897.85888671875,1433.4613037109375],"size":[158.98886108398438,46],"flags":{},"order":8,"mode":0,"inputs":[{"localized_name":"audio1","name":"audio1","type":"AUDIO","link":3},{"localized_name":"audio2","name":"audio2","type":"AUDIO","link":4}],"outputs":[{"localized_name":"audio_out","name":"audio_out","type":"AUDIO","links":[5]}],"properties":{"aux_id":"set-soft/ComfyUI-AudioBatch","ver":"20f2e30df280238fbbd1d4c19382140ff8684e21","Node name for S&R":"SET_AudioConcatenate"},"color":"#233","bgcolor":"#355"},{"id":8,"type":"MarkdownNote","pos":[1218.3682861328125,1269.0975341796875],"size":[251.47784423828125,116.92485046386719],"flags":{},"order":1,"mode":0,"inputs":[],"outputs":[],"properties":{},"widgets_values":["# a 3 seconds sweep"],"color":"#432","bgcolor":"#653"},{"id":9,"type":"MarkdownNote","pos":[1559.5836181640625,1190.826904296875],"size":[251.47784423828125,116.92485046386719],"flags":{},"order":2,"mode":0,"inputs":[],"outputs":[],"properties":{},"widgets_values":["# The first second"],"color":"#432","bgcolor":"#653"},{"id":10,"type":"MarkdownNote","pos":[1559.5836181640625,1642.667724609375],"size":[251.47784423828125,116.92485046386719],"flags":{},"order":3,"mode":0,"inputs":[],"outputs":[],"properties":{},"widgets_values":["# The last second"],"color":"#432","bgcolor":"#653"},{"id":11,"type":"MarkdownNote","pos":[2237.087158203125,1573.8836669921875],"size":[251.47784423828125,116.92485046386719],"flags":{},"order":4,"mode":0,"inputs":[],"outputs":[],"properties":{},"widgets_values":["# The first and last seconds"],"color":"#432","bgcolor":"#653"}],"links":[[1,1,0,2,0,"AUDIO"],[2,1,0,3,0,"AUDIO"],[3,2,0,4,0,"AUDIO"],[4,3,0,4,1,"AUDIO"],[5,4,0,5,0,"AUDIO"],[6,2,0,6,0,"AUDIO"],[7,3,0,7,0,"AUDIO"]],"groups":[],"config":{},"extra":{"ds":{"scale":0.8432169278238191,"offset":[-1105.7056732624972,-1037.8409022397152]}},"version":0.4} \ No newline at end of file diff --git a/source/nodes/nodes_audio.py b/source/nodes/nodes_audio.py index 3dc1119..0234e0e 100644 --- a/source/nodes/nodes_audio.py +++ b/source/nodes/nodes_audio.py @@ -774,3 +774,69 @@ class AudioSplit2Channels: } return (audio_left, audio_right) + + +class AudioConcatenate: + @classmethod + def INPUT_TYPES(cls): + return { + "required": { + "audio1": ("AUDIO", {"tooltip": "The first audio clip (or batch)."}), + "audio2": ("AUDIO", {"tooltip": "The second audio clip (or batch) to append."}), + }, + } + + RETURN_TYPES = ("AUDIO",) + RETURN_NAMES = ("audio_out",) + FUNCTION = "concatenate_audio" + CATEGORY = BASE_CATEGORY + "/" + MANIPULATION_CATEGORY + DESCRIPTION = "Concatenates two audio signals end-to-end in time." + UNIQUE_NAME = "SET_AudioConcatenate" + DISPLAY_NAME = "Audio Concatenate" + + def concatenate_audio(self, audio1: dict, audio2: dict): + # 1. Align the two audio inputs. This will unify their sample rate and channel count. + # It will also pad their *lengths* to be equal, which is not what we want for concatenation. + # We will use the aligned waveforms but ignore the padding by using their original lengths + # after resampling. + + aligner = AudioBatchAligner(audio1, audio2) + # aligned_wf1 is (B1, C_target, N_target), aligned_wf2 is (B2, C_target, N_target) + aligned_wf1, aligned_wf2, target_sr = aligner.get_aligned_waveforms() + + # 2. Determine the original lengths *after resampling* to know how much to take from each. + n1_orig = audio1['waveform'].shape[2] + n2_orig = audio2['waveform'].shape[2] + n1_after_resample = (n1_orig if audio1['sample_rate'] == target_sr else + int(n1_orig * (target_sr / audio1['sample_rate']))) + n2_after_resample = (n2_orig if audio2['sample_rate'] == target_sr else + int(n2_orig * (target_sr / audio2['sample_rate']))) + + # Take the un-padded, aligned data from each waveform + wf1_to_concat = aligned_wf1[..., :n1_after_resample] + wf2_to_concat = aligned_wf2[..., :n2_after_resample] + + # 3. Handle mismatched batch sizes by repeating the last item. + b1, b2 = wf1_to_concat.shape[0], wf2_to_concat.shape[0] + if b1 != b2: + if b1 < b2: + last_item = wf1_to_concat[-1:, :, :] # Keep batch dim + repeats_needed = b2 - b1 + wf1_to_concat = torch.cat([wf1_to_concat, last_item.repeat(repeats_needed, 1, 1)], dim=0) + else: # b2 < b1 + last_item = wf2_to_concat[-1:, :, :] + repeats_needed = b1 - b2 + wf2_to_concat = torch.cat([wf2_to_concat, last_item.repeat(repeats_needed, 1, 1)], dim=0) + + # Now both wf1_to_concat and wf2_to_concat are (max(B1,B2), C_target, N_relevant) + + # 4. Concatenate along the time/samples dimension (dim=2) + concatenated_waveform = torch.cat((wf1_to_concat, wf2_to_concat), dim=2) + + logger.info(f"Concatenated audio. Final shape: {concatenated_waveform.shape}, SR: {target_sr}") + + output_audio = { + "waveform": concatenated_waveform, + "sample_rate": target_sr + } + return (output_audio,) diff --git a/source/tests/test_audio_concatenate.py b/source/tests/test_audio_concatenate.py new file mode 100644 index 0000000..a78f854 --- /dev/null +++ b/source/tests/test_audio_concatenate.py @@ -0,0 +1,86 @@ +""" +Regression tests for the AudioConcatenate node in ComfyUI-AudioBatch. +""" + +import pytest +import torch + +import bootstrap # noqa: F401 +from nodes.nodes_audio import AudioConcatenate + + +# Helper function +def create_dummy_audio(batch_size, channels, samples, sr, device='cpu'): + waveform = torch.linspace(0.1, 0.9, samples).repeat(batch_size, channels, 1) + return {"waveform": waveform, "sample_rate": sr} + + +@pytest.fixture +def concat_node(): + return AudioConcatenate() + + +def test_concatenate_simple(concat_node): + """Tests concatenating two perfectly matched audio clips.""" + sr = 44100 + samples1, samples2 = 1000, 2000 + audio1 = create_dummy_audio(1, 2, samples1, sr) + audio2 = create_dummy_audio(1, 2, samples2, sr) + + (result_audio,) = concat_node.concatenate_audio(audio1, audio2) + + # Assertions + assert result_audio['sample_rate'] == sr + assert result_audio['waveform'].shape[0] == 1 # Batch size + assert result_audio['waveform'].shape[1] == 2 # Channels + assert result_audio['waveform'].shape[2] == samples1 + samples2 # Length + + # Check content + assert torch.allclose(result_audio['waveform'][..., :samples1], audio1['waveform']) + assert torch.allclose(result_audio['waveform'][..., samples1:], audio2['waveform']) + + +def test_concatenate_with_alignment(concat_node): + """Tests that alignment (SR, channels) happens correctly before concatenation.""" + # Audio1: Stereo, 44.1k SR, 1s long + sr1, len1 = 44100, 44100 + audio1 = create_dummy_audio(1, 2, len1, sr1) + + # Audio2: Mono, 22.05k SR, 1s long + sr2, len2 = 22050, 22050 + audio2 = create_dummy_audio(1, 1, len2, sr2) + + (result_audio,) = concat_node.concatenate_audio(audio1, audio2) + + # Expected result: + # SR = 44100 + # Channels = 2 + # Length = len1 + (len2 * sr1/sr2) = 44100 + (22050 * 2) = 88200 + expected_len = len1 + int(len2 * (sr1 / sr2)) + + assert result_audio['sample_rate'] == sr1 + assert result_audio['waveform'].shape == (1, 2, expected_len) + + +def test_concatenate_batch_mismatch(concat_node): + """Tests that batch mismatch is handled by repeating the last item.""" + sr = 44100 + samples = 1000 + # audio1 has 2 items, audio2 has 1 item + audio1 = create_dummy_audio(2, 1, samples, sr) + audio2 = create_dummy_audio(1, 1, samples, sr) + + (result_audio,) = concat_node.concatenate_audio(audio1, audio2) + + # Output batch should be max(2, 1) = 2 + assert result_audio['waveform'].shape[0] == 2 + assert result_audio['waveform'].shape[2] == samples * 2 # Concatenated length + + # The second item's second half should be a repeat of audio2's only item + # First, let's get the second half of the second batch item + second_item_second_half = result_audio['waveform'][1, :, samples:] + + # This should be equal to audio2's (only) item + audio2_item = audio2['waveform'][0, :, :] + + assert torch.allclose(second_item_second_half, audio2_item)